We provide IT Staff Augmentation Services!

Hadoop Developer Resume

0/5 (Submit Your Rating)

Aliso Viejo, CA

SUMMARY

  • Over 8 years of diversified IT experience in E2E data analytics platforms (ETL - BI-Java) as Big data, Hadoop, Java/J2EE Development and System Analysis.
  • Worked for over 4 years with Big Data/Hadoop Ecosystem in the implementation of Data Lake.
  • Hands on experience Hadoop framework and its ecosystem like Distributed file system (HDFS), MapReduce, Pig, Hive, Sqoop, Flume and Spark.
  • Experience in layers of Hadoop Framework - Storage (HDFS), Analysis (Pig and Hive), Engineering (Jobs and Workflows), extending the functionality by writing custom UDFs.
  • Extensive experience in developing Data warehouse applications using Hadoop, Informatica, Oracle, Teradata, MS SQL server on UNIX and Windows platforms and experience in creating complex mappings using various transformations and developing strategies for Extraction, Transformation and Loading (ETL) mechanism by using Informatica 9.x/8.x.
  • Proficient in Hive Query language and experienced in hive performance optimization using Static-Partitioning, Dynamic-Partitioning, Bucketing and Parallel Execution concepts.
  • As ETL developer, designed and maintained high performance ELT/ETL processes.
  • Experience in analyzing data using Hive QL, Pig Latin, and custom MapReduce programs in Java, custom UDF s.
  • Good Understanding of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and MapReduce concepts.
  • Knowledge on Cloud computing infrastructure AWS (amazon web services).
  • Created modules for spark streaming in data into Data Lake using Strom and Spark.
  • Experience in Dimensional Data Modeling Star Schema, Snow-Flake Schema, Fact and Dimensional Tables, concepts like Lambda Architecture, and Batch processing, Oozie.
  • Extensively used Informatica client tools Source Analyzer, Warehouse designer, Mapping designer, Mapplet Designer, ETL Transformations, Informatica Repository Manager and Informatica Server Manager, Workflow Manager & Workflow Monitor.
  • Expertise in using core Java, J2EE, Multithreading, JDBC, Shell Scripting and proficient in using Java API's Collections, Servlets, JSP for application development.
  • Worked closely to review pre- and post-processed data to ensure data accuracy and integrity with Dev and QA teams.
  • Experience in Java, J2ee, JDBC, Collections, Servlets, JSP, Struts, Spring, Hibernate, JSON, XML, REST, SOAP Web services, Groovy, MVC, Eclipse, Weblogic, Websphere, and Apache Tomcat severs.
  • Extensive knowledge of Data Modeling, Data Conversions, Data integration and Data Migration with specialization in Informatica Power Center.
  • Expertise in extraction, transformation and loading data from heterogeneous systems like flat files, excel, Oracle, Teradata, MSSQL Server.
  • Good work experience with UNIX/Linux commands, scripting and deploying the applications on the servers. Maintained tuning, and monitoring Hadoop jobs and clusters in a production environment.
  • Strong skills in algorithms, data structures, Object oriented design, Design patterns, documentation and QA/testing.
  • Excellent domain knowledge in Insurance, Telecom and Banking.

TECHNICAL SKILLS

BigData Technologies: Hortonworks HDP, Hadoop, MapReduce, Pig, Hive, Apache Spark, SQL, Informatica Power Center 9.6.1/8.x, Hbase/Cassandra, Kafka, Kibana, Storm, NoSQL, Elastic MapReduce(EMR), Tez, Impala, Hue, YARN, Mesos.

Databases: Hortonworks HDP, Oracle 10g/11g, Teradata, DB2, Microsoft SQL Server, MySQL, MongoDB, NoSQL, SQL databases.

Platforms (O/S): Red-Hat LINUX, Ubuntu, Windows NT/2000/XP.

Programming languages: Java, Scala, SQL, UNIX shell script, JDBC, Python, Perl.

Security Management: Hortonworks Ambari, Cloudera Manager, Apache Knox, XA Secure, Kerberos.

Data warehousing: Informatica Power center/Powermart/Data quality/Big data, Pentaho, ETL Development, Amazon Redshift, IDQ.

Database Tools: JDBC, HADOOP, Hive, No-SQL, SQL Navigator, SQL Developer, TOAD, SQL Plus, SAP Business Objects

Data Modeling: Rational Rose, Erwin 7.3/7.1/4.1/4.0

Editors: Eclipse, Intellij, Spark Eclipse

PROFESSIONAL EXPERIENCE

Confidential, Aliso Viejo, CA

Hadoop Developer

Responsibilities:

  • Gathered requirements and prepared document for the Big Data project and worked with Horton Works distribution.
  • Worked on BigData Analytics project to load the data from Source all through into Client's Modern Analytics Platform.
  • Analyzed and ingested Policy, Claims, Billing and Agency Data in Client's Solution, which is done through multiple stages.
  • Worked in an Agile SDLC to deliver new platform services and components.
  • Worked on Spark for in memory commutations and comparing the Data Frames for optimizing performance.
  • Worked on loading and transforming of large sets of data using Hive and Spark.
  • Implemented Kafka, spark streaming pipe lines to ingest real streaming data.
  • Developed MapReduce programs to process the Avro files and to get the results by performing some calculations on data and also performed map side joins. Supported MapReduce Java programs those are running on the cluster.
  • Imported Bulk Data into HBase Using MapReduce programs.
  • Used Rest API to Access HBase data to perform analytics.
  • Perform analytics on Time Series Data exists in HBase using HBase API.
  • Designed and implemented Incremental Imports into Hive tables.
  • Involved in creating Hive tables, loading with data and writing Hive queries that will run internally in MapReduce way.
  • Involved in collecting, aggregating and moving data from servers to HDFS using Flume.
  • Imported and Exported Data from Different Relational Data Sources like DB2,SQL Server, Teradata to HDFS using Sqoop.
  • Migrated complex map reduce programs into In-memory Spark processing using Transformations and actions.
  • Collected the real-time data from Kafka using Spark Streaming and performed transformations and aggregation on the fly to build the common learner data model and persists the data into Hbase.
  • Used SCALA to store streaming data to HDFS and to implement Spark for faster processing of data.
  • Worked on creating the RDD's, DF's for the required input data and performed the data transformations using SparkPython.
  • Involved in developing Spark SQLqueries, Data frames, import data from Data sources, perform transformations, and perform read/write operations, save the results to output directory into HDFS.
  • Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
  • Developed PIG scripts for the analysis of semi structured data.
  • Developed PIG UDF'S for manipulating the data according to Business Requirements and also worked on developing custom PIG Loaders.
  • Worked on Oozie workflow engine for job scheduling.
  • Developed Oozie workflow for scheduling and orchestrating the ETL process.
  • Managed and reviewed the Hadoop log files using Shell scripts.
  • Migrated ETL jobs to Pig scripts to do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
  • Worked on different file formats like Sequence files, XML files and Map files using MapReduce Programs.
  • Worked with Avro Data Serialization system to work with JSON data formats.
  • Used Amazon Web Services S3 to store large amount of data in identical/similar repository.
  • Worked with the Data Science team to gather requirements for various data mining projects.
  • Wrote shell scripts for rolling day-to-day processes and it is automated.
  • Involved in build applications using Maven and integrated with Continuous Integrationservers like Jenkins to build jobs.
  • Used Enterprise Data Warehouse database to store the information and to make it access all over organization.
  • Responsible for preparing technical specifications, analyzing functional Specs, development and maintenance of code.

Environment: Hadoop, Map Reduce, HDFS, Hive, Python, Scala, Kafka, Spark streaming, Spark Sql, MongoDB ETL, Oracle, Informatica 9.6,SQL, MapR, Sqoop, Zookeeper, AWS EMR,AWS S3,AWS EC2, Control-M scheduler, D3.JS,Jenkins, GIT, JIRA, Unix/Linux, Agile Methodology.

Confidential, Hoffman Estates, IL

Hadoop Developer

Responsibilities:

  • Worked on developing applications in Hadoop Big Data Technologies-Pig, Hive, MapReduce, Oozie, Flume, and Kafka.
  • Developed data pipeline using Flume, Sqoop, Pig and Java mapreduce to ingest customer behavioral data and financial histories into HDFS for analysis.
  • Used SQOOP, HDFS Put or Copy From Local to ingest data.
  • Used Pig to do transformations, event joins, filter bot traffic and some pre-aggregations before storing the data onto HDFS.
  • Developed Pig UDF’s for the needed functionality that is not out of the box available from Apache Pig.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
  • Developed Hive DDL’s to create, alter and drop Hive Tables.
  • Involved in developing Hive UDFs for the needed functionality that is not out of the box available from Apache Hive.
  • Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
  • Worked on Mapreduce Joins in querying multiple semi-structured data as per analytic needs.
  • Involved in loading data from Unix File System into HDFS with different format of data (Avro) and creating indexes and tuning the SQL queries in Hive and Involved in database connection by using Sqoop.
  • Involved in converting Hive/SQL queries into Spark functionality and analyze them using Scala API.
  • Worked on setting up High Availability for GPHD 2.2 with Zookeeper and quorum journal nodes.
  • Automated the process for extraction of data from warehouses and weblogs by developing work-flows and coordinator jobs in Oozie.
  • Worked in AWS environment for development and deployment of Custom HADOOP Applications.
  • Worked and learned a great deal from AmazonWebServices (AWS) Cloud services like EC2, S3 and EBS.
  • Used HCATALOG to access Hive table metadata from Pig code.
  • Responsible for developing data pipeline using flume, Sqoop and Pig to extract the data from weblogs and store in HDFS Designed and implemented various metrics that can statistically signify the success of the experiment.
  • Used Eclipse and ant to build the application.
  • Setup Hadoop cluster on Amazon EC2 using whirr for POC.
  • Involved in Loading process into the Hadoop distributed File System and Pig in order to preprocess the data.
  • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as MapReduce, Pig, Hive, Sqoop, flume) as well as some system specific jobs like shell scripts.
  • Imported and exported large sets of data into HDFS and vice-versa using Sqoop.
  • Used Sqoop for importing and exporting data into HDFS and Hive.
  • Processed ingested raw data using Map Reduce, Apache Pig and Hive.
  • Developed Pig Scripts for change data capture and delta record processing between newly arrived data and already existing data in HDFS.
  • Pivot the HDFS data from Rows to Columns and Columns to Rows.
  • Involved in emitting processed data from Hadoop to relational databases or external file systems using Sqoop, HDFS GET or Copy To Local.
  • Developed shell scripts to orchestrate execution of all other scripts (Pig, Hive, Map Reduce) and move the data files within and outside of HDFS.

Environment: Hadoop, Map Reduce, AWS, Spark, Scala, Kafka, Yarn, Hive, Pig, Hbase, Oozie, Sqoop, Flume, Oracle 11g, Core Java, Cloudera HDFS, Eclipse.

Confidential, Boston, MA

Hadoop Developer

Responsibilities:

  • Creating consolidated loss information file of various levels of business such as Claim, Policy, and Transaction and miscellaneous data.
  • Executed POCs for using Amazon Redshift, to test the feasibility of the DWH fit in our requirement.
  • Migration of the claims data in oracle to the analytical database created in Hadoop with Sqoop.
  • Worked extensively with Sqoop for importing metadata from Oracle.
  • This loss information file is supplied to mainframe for completing further business batch processes.
  • Key role in designing and implementing Mapreduce based applications for data validation. Data involves records and logs received from various production devices of the client.
  • Analyzed Session Log files in case the session fails in order to resolve errors in mapping or session configurations.
  • Worked with Data Governance team and implement the rules and build physical data model on hive in the data lake.
  • Exported the analysed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Analyze large and critical datasets using Cloudera, HDFS, Hbase, MapReduce, Hive, Hive UDF, Pig, Sqoop, and Zookeeper.
  • Used Pig to store the data into HBase, to parse the data and store in Avro format.
  • Created Hive tables, dynamic partitions, buckets for sampling, and worked on them using HiveQL.
  • Collected and aggregated large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
  • Worked with NoSQL databases like Hbase in creating Hbase tables to load large sets of semi structured data.
  • Mentored and delivered s to other team members on Hadoop ecosystem targeting MapReduce and Hive for cross-skill .
  • Written multiple MapReduce procedures to power data for extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV, Avro & other compressed file formats.
  • Performed unit testing, Data Reconciliation, knowledge transfer and mentored other team members.

Environment: Hadoop 2.0,Sqoop, Java, Apache Hbase, Informatica Power Center, IDQ analyst, DB Visualizer, Windows.

Confidential, Dublin, OH

Hadoop Developer

Responsibilities:

  • Involved in review of functional and non-functional requirements.
  • Installed and configured Hadoop Mapreduce, HDFS, Developed multiple Map Reduce jobs in java for data cleaning and pre-processing.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Supported Map Reduce Programs those are running on the cluster.
  • Involved in loading data from UNIX file system to HDFS.
  • Installed and configured Hive and written Hive UDFs.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Used Python and Django to interface with the jQuery UI and manage the storage and deletion of content.
  • Involved in Loading process into the Hadoop distributed File System and Pig in order to preprocess the data.
  • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as MapReduce, Pig, Hive, Sqoop) as well as some system specific jobs like (shell scripts)
  • Involved in Data modeling sessions to develop models for Hive tables
  • Imported and exported large sets of data into HDFS and vice-versa using Sqoop.
  • Transferred log files from the log generating servers into HDFS.
  • Worked on Integrating SSIS with Hadoop and performed ETL operations.
  • Created tasks, workflows and sessions using Workflow manager. Worked with scheduling team to come up Production schedule.
  • Worked on Hive partition and bucketing concepts and created hive External and Internal tables with Hive partition.
  • Used Ambari, Knox, Ranger to monitor the clusters.
  • Solved performance issues in Hive and pig with understanding of Joins, Group and aggregation and how does it transfer to Map-Reduce.
  • Moved the data from traditional databases like MySQL, MS SQL Server and Oracle into Hadoop.

Environment: Hadoop, Map Reduce, HDFS, Hive, Sqoop, UNIX Shell Scripting.

Confidential

Java Developer

Responsibilities:

  • Involving in Analysis, Design, Implementation and Bug Fixing Activities.
  • Involving in Functional & Technical Specification documents review.
  • Created and configured domains in production, development and testing environments using configuration wizard.
  • Involved in creating and configuring the clusters in production environment and deploying the applications on clusters.
  • Deployed and tested the application using Tomcat web server.
  • Analysis of the specifications provided by the clients.
  • Involved to Design of the Application.
  • Ability to understand Functional Requirements and Design Documents.
  • Developed Use Case Diagrams, Class Diagrams, Sequence Diagram, Data Flow Diagram
  • Coordinated with other functional consultants.
  • Web related development with JSP, AJAX, HTML, XML, XSLT, and CSS.
  • Create and enhance the stored procedures, PL/SQL, SQL for Oracle 9i RDBMS.
  • Designed and implemented a generic parser framework using SAX parser to parse XML documents which stores SQL.
  • Deployed the application on WebLogic Application Server 9.0.
  • Extensively used UNIX /FTP for shell Scripting and pulling the Logs from the Server.
  • Provided further Maintenance and support, this involves working with the Client and solving their problems which include major Bug fixing.

Environment: Java 1.4, Web logic Server 9.0, Oracle 10g, Web services Monitoring, Web Drive, UNIX/LINUX, Web Logic Server, JavaScript, HTML, CSS, XML

Confidential

JAVA- Designer and Developer

Responsibilities:

  • As a Software Engineer, involved in designing business layer and data management components using MVC frameworks such as Struts and Java/J2EE.
  • Requirement Analysis for the enhancements of the application.
  • Identify other source systems like Oracle, their connectivity, related tables and fields and ensure data integration of the job.
  • Preparation of project closure reports.
  • Writing the Junit test cases to all the components in the product.

Environment: Java, J2EE, Struts, JSP, Servlets, MS-SQL Server, Oracle-9i, Windows, UNIX.

Confidential

Designing and Developer

Responsibilities:

  • Requirement analysis.
  • Research and Development of this RTH management system.
  • Designing and developing the modules for writing test cases.

Technologies Used: Java, J2EE, Struts, MS-SQL Server.

Tools: Eclipse IDE.

Confidential

Developer and Team member

Responsibilities:

  • Requirement Analysis.
  • Designing, Coding and testing the application.
  • Writing the Unit Test cases.
  • Preparation of project reports.

Technologies Used: HTML, Google Maps, JavaScript, Java, J2EE, MySQL 5.0, MS4W 2.2.7

Tools: Edit plus, Eclipse.

We'd love your feedback!