We provide IT Staff Augmentation Services!

Hadoop Developer Resume

3.00/5 (Submit Your Rating)

IowA

PROFESSIONAL SUMMARY:

  • 7+ years of overall IT experience in a variety of industries, which includes hands on experience in Big Data technologies.
  • Expertise in writing Hadoop Jobs for analyzing structured and unstructured data using HDFS, Hive, HBase, Pig, Spark, Kafka, Scala, Oozie and Talend ETL.
  • Good knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN and MapReduce concepts.
  • Experience in working with different kind of MapReduce programs using Hadoop for working with Big Data analysis.
  • Experience in analyzing data using Hive QL, Pig Latin and custom MapReduce programs in Java
  • Experience in importing/exporting data using Sqoop into HDFS from Relational Database Systems and vice - versa.
  • Good at developing Big data based Solutions using Hadoop and Spark, Information Retrieval and Machine Learning areas.
  • Experienced in implementing Real-Time streaming and analytics using various technologies i.e. Spark Streaming and Kafka.
  • Hands-on experience in setting up Apache Hadoop and Cloudera CDH clusters on Ubuntu, Fedora and other Linux distributions environments.
  • Worked on NoSQL DB (Cassandra and HBase) for support enterprise production
  • Extensive experience in handling semi structured/unstructured data using Map Reduce Programs.
  • Experience in handling different file formats like text files, Sequence files, Avro data files, mahout xml files using Map Reduce programs.
  • Experienced with optimization techniques in sort and shuffle phase, compression techniques, jvm reuse, speculative execution, etc.
  • Experience in troubleshooting errors in HBase Shell/API, Pig, Hive and MapReduce.
  • Extensive experience working in Oracle, DB2, SQL Server and My SQL database.
  • Experienced in performing analytics on structured data using Hive queries, operations.
  • Worked on loading the production log files data into HDFS using FLUME.
  • Experience in reporting tools like Tableau and Spotfire.
  • Very good experience in complete project life cycle (design, development, testing and implementation).
  • Experience with Oozie Scheduler in setting up workflow jobs with Map/Reduce and Pig jobs
  • Good working experience using Sqoop to import data into HDFS from RDBMS.
  • Strong experience in database design, writing complex SQL queries and stored procedures using PL/SQL.
  • Experience in Java, JSP, Servlets, EJB, Hibernate, Spring, Java Script, Ajax, JQuery, XML, and HTML.

TECHNICAL SKILLS:

Big Data Technologies: Hadoop, MapReduce, HDFS, Hive, Pig, Zookeeper, Sqoop, Oozie, Flume, IMPALA, HBASE, Kafka, Storm.

Big Data Frameworks: HDFS, YARN, Spark.

Hadoop Distributions: Cloudera (CDH3, CDH4, CDH5), Horton works, Amazon EMR, EC2.

Programming Languages: Java, shell scripting, Scala.

Databases: RDBMS, MySQL, Oracle, Microsoft SQL Server, Teradata, DB2, PL/SQL, CASSANDRA, MongoDB.

Frameworks: Spring, Hibernate, JSF, EJB, JMS.

Operating System: Windows, Linux/Unix.

IDE and Tools: Eclipse, NetBeans, Tableau.

Scripting Languages: JSP & Servlets, JavaScript, XML, HTML, Python.

Application Servers: Apache Tomcat, Web Sphere, Web logic, JBoss.

Methodologies: Agile, SDLC, Waterfall.

Web Services: Restful, SOAP.

ETL Tools: Talend, Informatica.

Others: Solr, elastic search. 

PROFESSIONAL EXPERIENCE:

Confidential, Iowa

Hadoop Developer

Responsibilities:

  • Worked on analysing Hadoop stack and different big data analytic tools including Kafka, Pig and Hive, HBase database and Sqoop.
  • Designed high level ETL architecture for overall data transfer from the OLTP to OLAP.
  • Created various Documents such as Source-To-Target Data mapping Document, Unit Test Cases and Data Migration Document.
  • Worked on installing cluster, commissioning & decommissioning of Data nodes, Name node recovery, capacity planning, JVM tuning, map and slots configuration.
  • Developed Pig Latin scripts to extract the data from the web server output files to load in HDFS.
  • Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
  • Cluster co-ordination service through Zookeeper.
  • Created mappings using the transformations like Source Qualifier, Aggregator, Expression, Lookup, Router, Normalizer, Filter, Update Strategy and Joiner transformations.
  • Worked on Hive for exposing data for further analysis and for generating transforming files from different analytical formats to text files.
  • Implemented best income logic using Pig scripts and UDFs.
  • Designed and implemented Spark test bench application to evaluate quality of recommendations made by the engine. 
  • Tool monitored log input from several data centres, via Spark Stream, was analyse and data was parsed and saved into Cassandra.
  • Implemented Cluster balancing.
  • Migrated high-volume OLTP transactions from Oracle to Cassandra in order to reduce oracle licensing footprint.
  • Streaming and complex analyticsof processing are handled with use of Spark.
  • Implemented test scripts to support test driven development and continuous integration.
  • Worked on tuning the performance of Hive and Pig queries.
  • Worked on Impala for Massive parallel processing of Hive queries.
  • Streaming data to Hadoop using Kafka
  • Writing java code for custom partitioner and writable
  • Worked on the Analytics Infrastructure team to develop a stream filtering system on top of Apache Kafka
  • Worked on to ease the jobs by building the applications on top of NoSQL database Cassandra.
  • Configured Spark streaming to receive real time data from Kafka and store the stream data to HDFS
  • Unit tested and tuned SQLs and ETL Code for better performance.
  • Monitored the performance and identified performance bottlenecks in ETL code.
  • Used TABLEAU which grabs data to generate reports, graphs and charts summarising the given set of data
  • Worked on data utilizing a Hadoop, Zookeeper, and Accumulate stack, aiding in the development of specialized indexes for performant queries on big data implementations

Environment:Informatica Power Centre 9.5, Hadoop, HDFS, MapReduce, HBase, Hive, PIG, Sqoop, Oozie,Flume, Spark SQL, Spark Context, Spark Stream,Cassandra, Linux/Unix shell scripting, Big Data,Java, Tableau.

Confidential, Bolingbrook, IL

Hadoop Developer

Responsibilities:

  • Responsible for architecting Hadoop clusters with CDH3.
  • Extensively involved in Installation and configuration of Cloudera distribution Hadoop, NameNode, Job Tracker, Task Trackers and Data Nodes.
  • Commission or decommission the data nodes from cluster in case of problems.
  • Installed and configured Hadoop ecosystem like HBase, Flume, Pig and Sqoop.
  • Involved in Hadoop cluster task like Adding and Removing Nodes without any effect to running jobs and data.
  • Managed and reviewed Hadoop Log files.
  • Load log data into HDFS using Flume. Worked extensively in creating MapReduce jobs to power data for search and aggregation.
  • Creating mapping from Source to Target in Talend.
  • Worked extensively with Sqoop for importing metadata from Oracle.
  • Experience in data extraction into Data tax Cassandra cluster from Oracle (RDBMS) using Java Driver or Sqoop tools.
  • Responsible for smooth error-free configuration of DWH-ETL solution and Integration with Hadoop.
  • Designed a data warehouse using Hive.
  • Created partitioned tables in Hive.
  • Created user accounts and given users the access to the Hadoop cluster. 
  • Performed HDFS cluster support and maintenance tasks like adding and removing nodes without any effect to running nodes and data. 
  • Responsible for HBase REST server administration, backup and recovery .
  • Monitoring and controlling local file system disk space usage, log files, cleaning log files with automated scripts. 
  • As a Hadoop  admin, monitoring cluster health status on daily basis, tuning system performance related configuration parameters, backing up configuration xml files.
  • Monitored all MapReduce Read Jobs running on the cluster using Cloudera Manager and ensured that they were able to read the data to HDFS without any issues.
  • Involved in upgrading Hadoop Cluster from HDP 1.3 to HDP 2.0.
  • Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
  • Handle the upgrades and Patch updates.
  • Migrated ETL jobs to Pig scripts do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.

Environment: Hadoop, MapReduce, HDFS, Pig, Hive, HBase,Java,ETL,Oozie,Talend,Cloudera MySQL, Ubuntu,CDH3.

Confidential, San Diego, CA

Hadoop Developer

Responsibilities:

  • Mainly involved in creating and running Hadoop jobs to process raw binary data produced by vehicle sensors.
  • Massaging and parsing of the obtained raw binary data.
  • Gathering requirements from the product owner and the data science team.
  • Developed many MapReduce jobs in native Java for pre-processing of the data.
  • Developed Hive scripts to create both Internal and External Hive tables to store the transformed data.
  • Involved in creating and scheduling Oozie workflow scripts to run series of Sqoop imports, Mapreduce Transformation jobs, Hive scripts.
  • Created and maintained technical documentation for all the workflows.
  • Developed DFD (Data Flow Diagrams) on the company’s own Wiki for documentation and knowledge to understand the entire workflow of the project.
  • Worked with business analysts, Data science team and product owners to identify the tasks and obtain new requirements as part of agile scrum methodology.

Environment: MapReduce, Hive, Sqoop, Oozie, Hortonworks, HUE, Ambari, AVRO, Java, Hadoop, HDFS, Pig, and Big Data

Confidential

Hadoop Developer

Responsibilities:

  • Handle the installation and configuration of a Hadoop cluster.
  • Build and maintain scalable data pipelines using the Hadoop ecosystem and other open source components like Hive, and HBase. 
  • Handle the data exchange between HDFS and different web sources using Flume and Sqoop.
  • Monitor the data streaming between web sources and HDFS.
  • Monitor the Hadoop cluster functioning through monitoring tools.
  • Close monitoring and analysis of the MapReduce job executions on cluster at task level.
  • Inputs to development regarding the efficient utilization of resources like memory and CPU utilization based on the running statistics of Map and Reduce tasks
  • Changes to the configuration properties of the cluster based on volume of the data being processed and performance of the cluster.
  • Handle the upgrades and Patch updates.
  • Involved in loading data from UNIX file system to HDFS.
  • Installed and configured Hive and also written Hive QL scripts.
  • Responsible to manage data coming from different sources.
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
  • Extensive usage of Struts, HTML, CSS, JSP, JQuery, AJAX and JavaScript for interactive pages.
  • Used Ganglia to monitor the cluster around the clock.
  • Assist the team in their development & deployment activities.
  • Set up automated processes to analyze the System and Hadoop log files for predefined errors and send alerts to appropriate groups.
  • Set up and manage HA NameNode and NameNode federation using Apache 2.0 to avoid single point of failures in large clusters.
  • Set up the checkpoints to gathering the system statistics for critical set ups.
  • Discussions with other technical teams on regular basis regarding upgrades, Process changes, any Special processing and feedback.

Environment: Distcp,Kerberos,Ranger,HDFS,Pig,Hive,Sqoop,ZooKeeper,Oozie,Java,JEE,Spark,Solr,HBase.

Confidential

Java/J2EE developer

Responsibilities:

  • Involved in designing Class and Sequence diagrams with UML and Data flow diagrams.
  • Implemented MVC architecture using Struts framework to get the Free Quote.
  • Designed and developed front end using JSP, Struts (tiles), XML, JavaScript, and HTML.
  • Used Struts tag libraries to create JSP.
  • Implemented Spring MVC, dependency Injection (DI) and aspect oriented programming (AOP) features along with Hibernate.
  • Experienced with implementing navigation using Spring MVC.
  • Used Hibernate for object-relational mapping persistence.
  • Implemented message driven beans to get from queues to send again to support team using MSend commands.
  • Experienced with hibernate core interfaces like configuration, session factory, transactional and criteria interfaces.
  • Reviewed the requirements and Involved in database design for new requirements
  • Wrote Complex SQL queries to perform various database operations using TOAD.
  • Java Mail API was used to notify the Agents about the free quote and for sending Email to the Customer with Promotion Code for validation.
  • Involved in testing using Junit.
  • Performed application development using Eclipse and Web Sphere Application Server for deployment.
  • Used SVN for version control.

Environment:  Java, Spring, Hibernate, Jms, Web Services, Ejb, Sql/PlSql, Html, Css, Jsp, java script, Ant, Junit, Web sphere,  

Confidential

Java Developer

Responsibilities: 

  • Implemented server side programs by using Servlets and JSP 
  • Designed, developed and validated User Interface using HTML, Java Script, XML and CSS 
  • Implemented MVC using Struts Framework 
  • Implemented Controller Servlet to handle the access to database 
  • Participated in code walkthroughs, Debugging and defect fixing 
  • Involved in the co-ordination of end to end production release process 
  • Used SVN for versioning control system  
  • Developed PL/SQL stored procedures and triggers 
  • Used JDBC prepared statements to call from Servlets for database access 
  • Designed and documented the stored procedures 
  • Extensively used HTML for web based design 
  • Involved in writing JUnit Test Cases and done unit testing for various components 
  • Worked on database interaction layer for insertions, updating and retrieval operations of data from oracle database by writing stored procedures 
  • Used Spring Framework for Dependency Injection and integrated with Hibernate  
  • Used Log4J for any errors in the application 

Environment: Java, J2EE, JSP, Servlets, HTML, DHTML, XML, JavaScript, Struts, Eclipse, WebLogic, PL/SQL and Oracle

We'd love your feedback!