We provide IT Staff Augmentation Services!

Sr. Big Data Developer  Resume

4.00/5 (Submit Your Rating)

Memphis, Tn

SUMMARY:

  • IT professional with 8+ years of experience in Developing, implementing, testing and maintenance of various web applications using JAVA, J2EE technologies along with 4+ years of Big Data Analytics and Hadoop Ecosystems.
  • Excellent knowledge on Hadoop Architecture with Hadoop Distributed File System and expertise in various ecosystems Map Reduce, Pig, Hive, Hbase, Sqoop, Zookeeper, Oozie, Flume, Kafka, Apache Spark.
  • Excellent knowledge on Hadoop daemons and various components such as HDFS, Name Node, Data Node, Resource Manager, Node Manager and MapReduce programming paradigm.
  • Efficient in writing MapReduce Programs and using MapR API for analyzing the structured and unstructured data.
  • Experience in extending Pig and Hive core functionality by writing custom UDF’s for Data Analysis, Data transformation, file processing, and identifying user behavior by running Pig Latin Scripts.
  • Experience in installation, configuration, supporting and managing Hadoop Clusters using MapR, Cloudera, HortonWorks distributions.
  • Experienced in using Apache Flume for collecting, aggregation, moving large amount of data from application server to HDFS.
  • Experience in importing and exporting data between HDFS and RDBMS using Sqoop and vice - versa. Also, worked on incremental import by creating Sqoop metastore jobs.
  • Expert in working with Hive creating tables, implementing complex business logic, methods for segmenting large data sets by Partitioning and Bucketing to improve query performance
  • Good knowledge on No-SQL databases used for operational use cases like recommendation engine -Cassandra and Hbase.
  • Experienced in running query using Impala and used BI tools to run ad-hoc queries directly on Hadoop.
  • Experience in using Zookeeper for reliable distributed coordination and Oozie for scheduling workflows to manage Hadoop jobs.
  • Good Knowledge on setting Hadoop Cluster, Architecture and monitoring the cluster.
  • Good Knowledge in Amazon AWS concepts like EMR and EC2 web services which provides fast and efficient processing of Big Data.
  • Experience in working with BI team and transform Big Data requirements into Hadoop centric technologies.
  • Basic knowledge in application design using Unified Modeling Language (UML), Sequence diagrams, Case diagrams, Entity Relationship Diagrams (ERD) and Data Flow Diagrams
  • Exposure to Maven/Ant, Git along with Shell Scripting for Build & Deployment Process.
  • Extensive experience in migrating ETL operations into HDFS systems using Pig Scripts.
  • Used hive and pig to perform data transformation and aggregations.
  • Worked on File Formats like Parquet, TEXT, ORC for HIVE Querying and Processing.
  • Expertise in developing and implementing web applications in Core Java, J2EE (Servlets, JSP, JDBC, JSTL) and Web Services.
  • Experience in using version control management tools like CVS, SVN and Rational Clear Case.
  • Good understanding of IBM WebSphere Application Server, Apache Tomcat, JBoss and WebLogic in the areas of development, deployment, configuration settings and deployment descriptors.
  • Excellent communication, interpersonal, analytical skills, and strong ability to perform as part of team and keen in learning emerging technology.

TECHNICAL SKILLS:

Big Data Technologies: HDFS, MapReduce, Hive, Pig, Sqoop, Flume, Oozie, Hadoop Streaming, Zookeeper, Impala, Apache Spark, Kafka, SparkSQL

Hadoop Distributions: Cloudera, HortonWorks, MapR

Languages: Java, C, C++, SQL, PL/SQL, PIG-Latin, HQL, SCALA, Python

Web Technologies: HTML5, CSS3, JavaScript, JSP, JSF

Java & J2EE Technologies: Core Java, Servlets, JDBC, XML, Hibernate, Spring, Struts

Web Services: SOAP, REST, WSDL, JSON

Cloud Computing: AWS, Amazon EC2, S3

Application Servers: Jboss, Tomcat, Web Logic, Web Sphere

Reporting Tools /ETL Tools: Tableau, Talend

Databases: Oracle10g, MySQL, DB2, Derby, SQLServer

NOSQL Databases: Apache HBase, Cassandra

Version Control /Tracking Tools: RTC, Rational Clear case, Visual SourceSafe (VSS), SVN, CVS

PROFESSIONAL EXPERIENCE:

Confidential, Memphis, TN

Sr. Big Data Developer

Responsibilities:

  • Developed Map reduce program which were used to extract and transform the data sets and result dataset were loaded to Cassandra and vice versa using kafka.
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Exploring with the Spark, improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark -SQL, Data Frame, Spark YARN.
  • Batch processing of data sources using Apache Spark, Elastic search.
  • Developed analytical components using Scala, Spark, Apache Mesos and Spark Stream.
  • Import the data from different sources like HDFS/Hbase into Spark RDD.
  • Converting Map Reduce programs into Spark transformations using Spark RDD's on Scala.
  • Implemented spark advanced procedures like text analytics and processing using the in-memory computing capabilities.
  • Developed Spark scripts by using Scala Shell commands as per the requirement.
  • Used different data Modeling and Data Warehouse design and development.
  • Developed numerous MapReduce jobs in java for Data Cleansing and analyzing data in PIG and Hive.
  • Bulk importing of data from various data sources, transform data in flexible ways by using flume, Map Reduce.
  • Loaded and Extracted the data using Sqoop from Oracle into HDFS.
  • Leverage Pig, Hive, Impala to analyze large data sets on Hadoop to study customer behavior.
  • Developed Pig Latin scripts for cleansing the data, precomputing common aggregates before loading into data warehouse.
  • Worked on implementation and maintenance of Cloudera Hadoop cluster.
  • Keeping a check on system health, logs and react accordingly to failure conditions.
  • Produced a quality Technical documentation for operating, complex configuration management, architecture changes and maintaining HADOOP Clusters.

Environment: Hadoop, HDFS, Spark, MapReduce, Pig, Hive, Impala, Sqoop, Kafka, HBase, Oozie, Flume, Scala, Python, Java, SQL Scripting and Linux Shell Scripting, Cloudera, EC2, EMR, S3, Oracle.

Confidential, San Francisco, CA

Hadoop Developer

Responsibilities:

  • Worked on writing complex Map-Reduce code in Java for processing data in multiple clusters.
  • Creating Hive Tables using HiveQL, loading data and running Hive queries to invoke underlying MapReduce program.
  • Involved in loading data into HBase using HBase Shell, HBase Client API, Pig and Sqoop.
  • Incrementally Updating of Hive Tables Using Sqoop.
  • Setting up and Deploying Apache SOLR to provide significantly better search performance.
  • Installing, configuring, supporting and managing Hadoop Clusters using Apache Cloudera.
  • Involved in using Ingestion mechanism tool Flume for collecting, aggregating and transporting streaming data from various webservers to a centralized data store HDFS.
  • Optimizing Hive queries on Data layout techniques such as partitioning and bucketing, Data sampling, Data processing and using custom file formats.
  • Experienced in managing and reviewing the Hadoop log files.
  • Built a Pig application of ETL transaction model by ingesting data from files, streams using the UDFs to perform select, iteration and other transforms over the data and finally store the results into the Hadoop Data File System.
  • Creating workflows and managing coordination among jobs using Oozie and automate tasks.
  • Worked with Avro Data Serialization system to work with JSON data formats.
  • Worked on different file formats like Sequence files, XML files and Map files using Map Reduce Programs.

Environment: Hadoop, Big Data, HDFS, MapReduce, Sqoop, Oozie, Pig, Hive, Hbase, Flume, LINUX, Java, Eclipse, Hue, Cassandra, Cloudera Manager, PL/SQL, Toad 9.6, UNIX Shell Scripting, Putty and Eclipse.

Confidential, Bridgewater, NJ

Hadoop Developer

Responsibilities:

  • Imported Data from Different Relational Data Sources like RDBMS, Teradata to HDFS using Sqoop.
  • Imported Bulk Data into HBase Using Map Reduce programs.
  • Perform data fetching and analytics on Time Series Database using HBase API.
  • Designed and implemented Incremental Imports into Hive tables.
  • Used Rest API to query HBase data and perform analytics.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS& Extracted the data from MySQL into HDFS using Sqoop.
  • Worked in Loading and transforming large sets of structured, semi structured and unstructured data
  • Involved in collecting, aggregating and moving data from servers to HDFS using Apache Flume.
  • Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
  • Using Sqoop to export data into HDFS and Hive and create reports for BI team.
  • Involved in creating Hive tables, and loading them into dynamic partition tables.
  • Experienced in managing and reviewing the Hadoop log files.
  • Using Pig to perform data operations ETL data pipelines, research on raw data and iterative data processing before storing the data to HDFS.
  • Worked on different file formats like Sequence files, XML files and Map files using Map Reduce Programs.
  • Developed multiple MapReduce jobs in Java for data cleansing and preprocessing.
  • Importing data from RDBMS to HDFS using Sqoop for generating reports and visualizations using Tableau.
  • Worked on Oozie workflow engine for job scheduling.

Environment: Hadoop, HDFS, Map Reduce, Hive, HBase, Oozie, Sqoop, Pig, Java, Tableau, Rest API, Maven.

Confidential

Sr. Java Developer / Hadoop Developer

Responsibilities:

  • Involved in planning, designing, developing and testing high quality software by Software Development Life Cycle (SDLC) process.
  • Assisted the analysis team in performing the feasibility analysis of the project.
  • Use Case diagrams, Class diagrams, Sequence diagrams and Object Diagrams were designed using Rational Rose 4.0.
  • Presentation Layer of the project was developed using JSF, Servlet technologies and UI technologies like HTML, CSS and JavaScript.
  • Developed Entity Java Beans (EJB) classes to implement various business functionalities (session beans).
  • Business tier was developed using Stateless and Stateful Session beans with EJB 2.0 standards using Web Sphere Studio Application Developer (WSAD 5.0).
  • Used Java API Collections like Lists, Sets and Maps.
  • Used SQL statements, stored procedures, handled SQL Injections and persisted data using Hibernate Sessions, Transactions and SessionFactoryObjects.
  • Integrated Spring DAO was used for data access using Hibernate.
  • Hibernate mapping files were created to map entity beans to tables in Database.
  • Ant build scripts were used for compiling and building the project.
  • Used IBM Websphere portal and IBM Websphere application server for deploying the applications.
  • CVS Repository was used for Version Control.
  • Created test plans and JUnit and Test Suite were used for testing the application.
  • Good hands on experience in UNIX commands used to see the log files on the server.
  • Assisted in developing Testing Plans and procedures for unit test, system test.
  • As part of development Unit Test Case Preparation and Unit Testing was done.

Environment: Java 1.4.1, JSP 2.0, HTML5, JavaScript, EJB 2.0, Struts 1.1, JDBC 2.0, IBM Web Sphere 5.0, XML, XSLT, XML Schema, JUnit 3.8.1, Rational Rose 4.0, Ant 1.5, UML, Hibernate 3, Linux, Oracle 9i and Windows.

Confidential

Java Developer

Responsibilities:

  • Technical responsibilities included high level architecture and rapid development.
  • Design architecture following J2EE MVC framework.
  • Inspection/Review of quality deliverables such as Design Documents.
  • Created UML class diagrams that depict the code's design and its compliance with the functional requirements.
  • Involved in designing & developing web-services using SOAP and WSDL.
  • Designed and developed the UI using Struts view component, JSP, HTML, CSS and JavaScript.
  • Developed and implemented Servlets running under JBoss.
  • Used J2EE design patterns and Data Access Object (DAO) for the Business Tier and Persistence Tier layer of the project.
  • Developed various EJBs for handling business logic and data manipulations in database.
  • Mapped CMP entity beans to database based on business logic.
  • Development of database interaction code to JDBC API making extensive use of SQL Query Statements and advanced prepared statement.
  • Wrote SQL Scripts, Stored procedures and SQL Loader to load reference data.
  • Involved in writing Spring Configuration XML files that contains declarations and other dependent object declarations.
  • Involved in creation and running of Test Cases for JUnit Testing.
  • Log4J components for logging was used. Performed daily monitoring of log files and resolved issues.
  • Experience in implementing Web Services using SOAP, REST and XML/HTTP technologies.

Environment: Java, J2EE, Spring, JSP, Hibernate, Java Script, CSS, JDBC, IntelliJ, LDAP, REST, Active Directory, SAML, Web Services, Microsoft SQL Server, HTML.

We'd love your feedback!