We provide IT Staff Augmentation Services!

Hadoop Developer Resume

5.00/5 (Submit Your Rating)

Sacramento, CA

PROFESSIONAL SUMMARY:

  • 6+ years of experience in the IT industry, which includes experience in Hadoop and Java application development
  • Hands on experience in Hadoop ecosystem components like Map Reduce, HDFS, Sqoop, Pig, Hive and Oozie.
  • Extensively working on Spark and Shark.
  • Working on Spark Streaming with Flume online streaming.
  • Good Experience on Media Analytics.
  • Expert in working with Hive data warehouse tool - creating tables, data distribution by implementing partitioning and bucketing, writing and optimizing the HiveQL queries.
  • Experience in using Apache Sqoop to import and export data to and from HDFS and Hive.
  • Hands on experience in setting up workflow using Apache Oozie workflow engine for managing and scheduling Hadoop jobs.
  • Experience in NoSQL Column-Oriented Databases like HBase, Cassandra, Mongo DB and its Integration with Hadoop cluster.
  • Experience in using Hcatalog for Hive, Pig and HBase and also has experience in troubleshooting errors in HBase Shell, Hive, Pig, Map reduce.
  • Experience in working with BI team and transform Big Data requirements into Hadoop centric technologies.
  • Experience in Hadoop Map Reduce programming, Pig scripting, HiveQL and HDFS.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems.
  • Experience with Oozie workflow engine to run multiple Hive and Pig jobs independently with time and data availability.
  • Strong experience on Hadoop distributions Horton works and Cloudera
  • Experience in data management and implementation of Big Data applications using Hadoop Frameworks.
  • Excellent understanding and knowledge of job workflow scheduling and locking tools/services like Oozie and Zookeeper.
  • Experience in writing Map Reduce jobs using Java and executing the jobs and also troubleshooting them.
  • Built real-time Big Data solutions using HBASE handling billions of records.
  • Extensive knowledge and work Experience in Systems Analysis, Design, Development, Implementation and Testing of Application software for Business solutions, Database Management, Data Analytics.
  • Worked on Classic and Yarn distributions of Hadoop like the Apache Hadoop 2.0.0.
  • Experience in data extraction and transformation using MapReduce jobs.
  • Having Good Knowledge on single node and multinode cluster configurations.
  • Experience in application development using Core Java, J2EE, Hibernate, JDBC, JSP and Servlets.
  • Good knowledge in Java Swing, JUnit, CSS, HTML, Java Applets, Apache Ant.
  • Proficient in database development: Oracle, PL SQL.
  • Good understanding of some of key design concepts, design patterns, and UML.
  • Experience in all phases of systems development.
  • Strong technical and interpersonal skills combined with great commitment towards meeting deadlines.
  • Experience working in both team and individual environments. Always eager to learn new technologies and implement them in challenging environment.
  • Excellent written and verbal communication skills.
  • Strong analytical and problem solving skills.
  • Excellent problem solving skills and understanding skills

TECHNICAL SKILLS:

Big Data Ecosystems: Hadoop, MapReduce, HDFS, HBase, Hive, Pig, Sqoop, Spark, Storm,Kafka, Oozie

Languages: C, Core Java, Unix, SQL,Python, R, C#, Haskell, Scala

J2EE Technologies: Servlets, JSP, JDBC, Java Beans

Methodologies: Agile, UML, Design Patterns (Core Java and J2EE

NoSQL Technologies: Cassandra, Mongo DB, Neo4j, HBase

Frameworks: MVC, Struts, Hibernate, Spring

Database: Oracle 11g, MySQL, MS-SQL Server, Teradata. PostgreSQL, IBM DB2

Operating Systems: Windows XP/Vista/7, UNIX

Development / Build Tools: Eclipse, Ant, Maven

Software Package: MS Office 2010

Tools: & Utilities: Eclipse, Net Beans, My Eclipse, SVN, Git, Maven, SOAP UI, JMX explorer, XML Spy, QC, QTP, Jira

Web Servers: WebLogic, WebSphere, Apache Tomcat.

Web Technologies: HTML,XML,JavaScript, jQuery, AJAX, SOAP, and WSDL.

WORK EXPERIENCE:

Confidential, Sacramento, CA

Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Written multiple MapReduce programs in Java for Data Analysis
  • Wrote MapReduce job using Pig Latin and Java API
  • Performed performance tuning and troubleshooting of MapReduce jobs by analyzing and reviewing Hadoop log files.
  • Developed pig scripts for analyzing large data sets in the HDFS.
  • Collected the logs from the physical machines and the Open Stack controller and integrated into HDFS using Flume
  • Experienced in migrating Hive QL into Impala to minimize query response time.
  • Knowledge on handling Hive queries using Spark SQL that integrate with Spark environment.
  • Implemented Avro and parquet data formats for apache Hive computations to handle custom business requirements.
  • Responsible for creating Hive tables, loading the structured data resulted from MapReduce jobs into the tables and writing hive queries to further analyze the logs to identify issues and behavioral patterns.
  • Worked on Sequence files, RC files, Map side joins, bucketing, partitioning for Hive performance enhancement and storage improvement.
  • Worked on analyzing Hadoop cluster using various Big Data ecosystems including Hive, Sqoop, Pig, Flume, HBase
  • Importing the data from Oracle into the HDFS using Sqoop. Performed full and incremental imports using Sqoop jobs.
  • Responsible to manage data coming from various sources and involved in HDFS maintenance and loading of structured and unstructured data.
  • Used Hive to form an abstraction on top of structured data that resides in HDFS and implemented Partitions, Dynamic Partitions, Buckets on HIVE tables.
  • Performed extensive Data Mining applications using HIVE.
  • Implemented Daily Cron jobs that automate parallel tasks of loading the data into HDFS using autosys and Oozie coordinator jobs.
  • Performed streaming of data into Apache ignite by setting up cache for efficient data analysis.
  • Responsible for performing extensive data validation using Hive
  • Sqoop jobs, PIG and Hive scripts were created for data ingestion from relational databases to compare with historical data.
  • Involved in submitting and tracking MapReduce jobs using Job Tracker.
  • Involved in creating Oozie workflow and Coordinator jobs to kick off the jobs on time for data availability.
  • Used Pig as ETL tool to do transformations, event joins, filter and some pre-aggregations
  • Responsible for cleansing the data from source systems using Ab Initio components such as Join, Dedup Sorted, De normalize, Normalize, Reformat, Filter-by-Expression, Rollup.
  • Used Visualization tools such as Power view for excel, Tableau for visualizing and generating reports.
  • Exported data to Tableau and excel with Power view for presentation and refining
  • Implemented business logic by writing Pig UDFs in Java and used various UDFs from Piggy Banks and other sources.
  • Implemented Hive Generic UDFs to implement business logic.
  • Coordinated with end users for designing and implementation of analytics solutions for User Based Recommendations using as per project proposals.
  • Implemented test scripts to support test driven development and continuous integration.
  • Involved in story-driven agile development methodology and actively participated in daily scrum meetings.

Environment: Hadoop, MapReduce, HDFS, Pig, Hive, Sqoop, Flume, Oozie, Java, Linux, Maven, Teradata, Zookeeper, SVN, autosys, Tableau, HBase, Cassandra, Mongo DB, Apache ignite.

Confidential, Houston, Texas

Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data pipelines using Hadoop.
  • Wrote Pig scripts to debug Kafka hourly data and perform daily roll ups.
  • Data Migration from existing Teradata systems to HDFS and build datasets on top of it.
  • Built a framework using SHELL scripts to automate Hive registration, which does dynamic table creation and automated way to add new partitions to the table.
  • Designed Hive external tables using shared meta-store instead of derby with partitioning, dynamic partitioning and buckets.
  • Involved in managing nodes on Hadoop cluster and monitor Hadoop cluster job performance using Cloudera manager.
  • Developed optimal strategies for distributing the web log data over the cluster importing and exporting the stored web log data into HDFS and Hive using Sqoop.
  • Involved in loading data from edge node to HDFS using shell scripting.
  • Created MapReduce programs to handle semi/unstructured data like xml, json, Avro data files and sequence files for log files.
  • Setup and benchmarked Hadoop/HBase clusters for internal use. Developed Simple to complex MapReduce programs.
  • Created and maintained Technical documentation for launching Cloudera Hadoop Clusters and for executing Hive queries and Pig Scripts.
  • Developed workflow-using Oozie for running MapReduce jobs and Hive Queries.
  • Implementing various advanced join operations using Pig Latin.
  • Done the work in importing and exporting data into HDFS and assisted in exporting analyzed data to RDBMS using SQOOP.
  • Assisted in exporting analyzed data to relational databases using Sqoop.
  • Involved in Develop monitoring and performance metrics for Hadoop clusters.
  • Worked with both MapReduce 1 (Job Tracker) and MapReduce 2 (YARN).
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
  • Optimized MapReduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Developed Oozie workflows that chain Hive/MapReduce modules for ingesting periodic/hourly input data.
  • Wrote Pig & Hive scripts to analyze the data and detect user patterns.
  • Implemented Device based business logic using Hive UDFs to perform ad-hoc queries on structured data.
  • Storing and loading the data from HDFS to Amazon S3 and backing up the Namespace data into NFS Filers.
  • Prepared Avro schema files for generating Hive tables and shell scripts for executing Hadoop commands for single execution.
  • Continuously monitored and managed the Hadoop cluster by using Cloudera Manager.
  • Worked with administration team to install operating system, Hadoop updates, patches, version upgrades as required.
  • Developed ETL pipelines to source data to Business intelligence teams to build visualizations.
  • Involved in unit testing, interface testing, system testing and user acceptance testing of the workflow Tool.

Environment: Cloudera Manager, MapReduce, HDFS, Pig, Hive, Sqoop, Apache Kafka, Oozie, Teradata, Avro, Java (JDK 1.6), Eclipse.

Confidential, Charlotte, NC

Hadoop Developer

Responsibilities:

  • Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables.
  • Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with reference tables and historical metrics.
  • Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
  • Involved in Development and Implementation of business Applications using Java/J2EE Technologies.
  • Use of build script using ANT to generate JAR, WAR, EAR files and for integration testing and unit testing.
  • Developed the entire application implementing MVC Architecture integrating JSP with Hibernate and spring frameworks.
  • Created dynamic HTML pages, used JavaScript for client-side validations, and AJAX to create interactive front-end GUI.
  • Used J2EE Design/Enterprise Integration patterns and SOA compliance for design and development of applications.
  • Implemented AJAX functionality using jQuery and JSON to communicate to the server and populate the data on the JSP.
  • Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems.
  • Managed and reviewed Hadoop log files.
  • Shared responsibility for administration of Hadoop, Hive and Pig.

Environment: Hadoop 1x, Hive, Pig, HBASE, Sqoop and Flume, Spring, jQuery, Java, J2EE, HTML, JavaScript, Hibernate.

Confidential

Java Developer

Responsibilities:

  • Involved in the design and development phases of Rational Unified Process (RUP).
  • Developed JSP Pages made them accessible to the Client using JBoss Application Server.
  • Extensively used complex SQL statements including joins and nested queries
  • Involved in various Software Development Life Cycle (SDLC) phases of the project like Development, Enhancements and Maintenance.
  • Followed Struts MVC framework to develop the application.
  • Designed and implemented the UI using Java, HTML, JSP and JavaScript.
  • Designed and developed web pages using Servlets and JSPs and also used XML/XSL/XSLT as repository.
  • Involved in Java application testing and maintenance in development and production.
  • Involved in developing the customer form data tables. Maintaining the customer support and customer data from database tables in MySQL database.
  • Designed and developed Views, Model and Controller components implementing MVC Framework.
  • Coded JSP pages and used JavaScript for client side validations and to achieve other client-side functionality.
  • Developed Java Helper classes for updating Customer Accounts and Customer information.

Environment: Java, Eclipse, JDBC, Servlets, JSP, WebLogic Server, Struts MVC, HTML, CSS and MySQL.

We'd love your feedback!