Hadoop Developer Resume
IowA
PROFESSIONAL SUMMARY:
- 7+ years of overall IT experience in a variety of industries, which includes hands on experience in Big Data technologies.
- Expertise in writing Hadoop Jobs for analyzing structured and unstructured data using HDFS, Hive, HBase, Pig, Spark, Kafka, Scala, Oozie and Talend ETL.
- Good knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN and MapReduce concepts.
- Experience in working with different kind of MapReduce programs using Hadoop for working with Big Data analysis.
- Experience in analyzing data using Hive QL, Pig Latin and custom MapReduce programs in Java.
- Experience in importing/exporting data using Sqoop into HDFS from Relational Database Systems and vice - versa.
- Good at developing Big data based Solutions using Hadoop and Spark, Information Retrieval and Machine Learning areas.
- Experienced in implementing Real-Time streaming and analytics using various technologies i.e. Spark Streaming and Kafka.
- Hands-on experience in setting up Apache Hadoop and Cloudera CDH clusters on Ubuntu, Fedora and other Linux distributions environments.
- Worked on NoSQL DB (Cassandra and HBase) for support enterprise production
- Extensive experience in handling semi structured/unstructured data using Map Reduce Programs.
- Experience in handling different file formats like text files, Sequence files, Avro data files, mahout xml files using Map Reduce programs.
- Experienced with optimization techniques in sort and shuffle phase, compression techniques, jvm reuse, speculative execution, etc.
- Experience in troubleshooting errors in HBase Shell/API, Pig, Hive and MapReduce.
- Extensive experience working in Oracle, DB2, SQL Server and My SQL database.
- Experienced in performing analytics on structured data using Hive queries, operations.
- Worked on loading the production log files data into HDFS using FLUME.
- Experience in reporting tools like Tableau and Spotfire.
- Very good experience in complete project life cycle (design, development, testing and implementation).
- Experience with Oozie Scheduler in setting up workflow jobs with Map/Reduce and Pig jobs
- Good working experience using Sqoop to import data into HDFS from RDBMS.
- Strong experience in database design, writing complex SQL queries and stored procedures using PL/SQL.
- Experience in Java, JSP, Servlets, EJB, Hibernate, Spring, Java Script, Ajax, JQuery, XML, and HTML.
TECHNICAL SKILLS:
Big Data Technologies: Hadoop, MapReduce, HDFS, Hive, Pig, Zookeeper, Sqoop, Oozie, Flume, IMPALA, HBASE, Kafka, Storm.
Big Data Frameworks: HDFS, YARN, Spark.
Hadoop Distributions: Cloudera (CDH3, CDH4, CDH5), Horton works, Amazon EMR, EC2.
Programming Languages: Java, shell scripting, Scala.
Databases: RDBMS, MySQL, Oracle, Microsoft SQL Server, Teradata, DB2, PL/SQL, CASSANDRA, MongoDB.
Frameworks: Spring, Hibernate, JSF, EJB, JMS.
Operating System: Windows, Linux/Unix.
IDE and Tools: Eclipse, NetBeans, Tableau.
Scripting Languages: JSP & Servlets, JavaScript, XML, HTML, Python.
Application Servers: Apache Tomcat, Web Sphere, Web logic, JBoss.
Methodologies: Agile, SDLC, Waterfall.
Web Services: Restful, SOAP.
ETL Tools: Talend, Informatica.
Others: Solr, elastic search.
PROFESSIONAL EXPERIENCE:
Confidential, Iowa
Hadoop Developer
Responsibilities:
- Worked on analysing Hadoop stack and different big data analytic tools including Kafka, Pig and Hive, HBase database and Sqoop.
- Designed high level ETL architecture for overall data transfer from the OLTP to OLAP.
- Created various Documents such as Source-To-Target Data mapping Document, Unit Test Cases and Data Migration Document.
- Worked on installing cluster, commissioning & decommissioning of Data nodes, Name node recovery, capacity planning, JVM tuning, map and slots configuration.
- Developed Pig Latin scripts to extract the data from the web server output files to load in HDFS.
- Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
- Cluster co-ordination service through Zookeeper.
- Created mappings using the transformations like Source Qualifier, Aggregator, Expression, Lookup, Router, Normalizer, Filter, Update Strategy and Joiner transformations.
- Worked on Hive for exposing data for further analysis and for generating transforming files from different analytical formats to text files.
- Implemented best income logic using Pig scripts and UDFs.
- Designed and implemented Spark test bench application to evaluate quality of recommendations made by the engine.
- Tool monitored log input from several data centres, via Spark Stream, was analyse and data was parsed and saved into Cassandra.
- Implemented Cluster balancing.
- Migrated high-volume OLTP transactions from Oracle to Cassandra in order to reduce oracle licensing footprint.
- Streaming and complex analyticsof processing are handled with use of Spark.
- Implemented test scripts to support test driven development and continuous integration.
- Worked on tuning the performance of Hive and Pig queries.
- Worked on Impala for Massive parallel processing of Hive queries.
- Streaming data to Hadoop using Kafka.
- Writing java code for custom partitioner and writable
- Worked on the Analytics Infrastructure team to develop a stream filtering system on top of Apache Kafka.
- Worked on to ease the jobs by building the applications on top of NoSQL database Cassandra.
- Configured Spark streaming to receive real time data from Kafka and store the stream data to HDFS
- Unit tested and tuned SQLs and ETL Code for better performance.
- Monitored the performance and identified performance bottlenecks in ETL code.
- Used TABLEAU which grabs data to generate reports, graphs and charts summarising the given set of data
- Worked on data utilizing a Hadoop, Zookeeper, and Accumulate stack, aiding in the development of specialized indexes for performant queries on big data implementations
Environment:Informatica Power Centre 9.5, Hadoop, HDFS, MapReduce, HBase, Hive, PIG, Sqoop, Oozie,Flume, Spark SQL, Spark Context, Spark Stream,Cassandra, Linux/Unix shell scripting, Big Data,Java, Tableau.
Confidential, Bolingbrook, IL
Hadoop Developer
Responsibilities:
- Responsible for architecting Hadoop clusters with CDH3.
- Extensively involved in Installation and configuration of Cloudera distribution Hadoop, NameNode, Job Tracker, Task Trackers and Data Nodes.
- Commission or decommission the data nodes from cluster in case of problems.
- Installed and configured Hadoop ecosystem like HBase, Flume, Pig and Sqoop.
- Involved in Hadoop cluster task like Adding and Removing Nodes without any effect to running jobs and data.
- Managed and reviewed Hadoop Log files.
- Load log data into HDFS using Flume. Worked extensively in creating MapReduce jobs to power data for search and aggregation.
- Creating mapping from Source to Target in Talend.
- Worked extensively with Sqoop for importing metadata from Oracle.
- Experience in data extraction into Data tax Cassandra cluster from Oracle (RDBMS) using Java Driver or Sqoop tools.
- Responsible for smooth error-free configuration of DWH-ETL solution and Integration with Hadoop.
- Designed a data warehouse using Hive.
- Created partitioned tables in Hive.
- Created user accounts and given users the access to the Hadoop cluster.
- Performed HDFS cluster support and maintenance tasks like adding and removing nodes without any effect to running nodes and data.
- Responsible for HBase REST server administration, backup and recovery .
- Monitoring and controlling local file system disk space usage, log files, cleaning log files with automated scripts.
- As a Hadoop admin, monitoring cluster health status on daily basis, tuning system performance related configuration parameters, backing up configuration xml files.
- Monitored all MapReduce Read Jobs running on the cluster using Cloudera Manager and ensured that they were able to read the data to HDFS without any issues.
- Involved in upgrading Hadoop Cluster from HDP 1.3 to HDP 2.0.
- Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
- Handle the upgrades and Patch updates.
- Migrated ETL jobs to Pig scripts do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
Environment: Hadoop, MapReduce, HDFS, Pig, Hive, HBase,Java,ETL,Oozie,Talend,Cloudera MySQL, Ubuntu,CDH3.
Confidential, San Diego, CA
Hadoop Developer
Responsibilities:
- Mainly involved in creating and running Hadoop jobs to process raw binary data produced by vehicle sensors.
- Massaging and parsing of the obtained raw binary data.
- Gathering requirements from the product owner and the data science team.
- Developed many MapReduce jobs in native Java for pre-processing of the data.
- Developed Hive scripts to create both Internal and External Hive tables to store the transformed data.
- Involved in creating and scheduling Oozie workflow scripts to run series of Sqoop imports, Mapreduce Transformation jobs, Hive scripts.
- Created and maintained technical documentation for all the workflows.
- Developed DFD (Data Flow Diagrams) on the company’s own Wiki for documentation and knowledge to understand the entire workflow of the project.
- Worked with business analysts, Data science team and product owners to identify the tasks and obtain new requirements as part of agile scrum methodology.
Environment: MapReduce, Hive, Sqoop, Oozie, Hortonworks, HUE, Ambari, AVRO, Java, Hadoop, HDFS, Pig, and Big Data
Confidential
Hadoop Developer
Responsibilities:
- Handle the installation and configuration of a Hadoop cluster.
- Build and maintain scalable data pipelines using the Hadoop ecosystem and other open source components like Hive, and HBase.
- Handle the data exchange between HDFS and different web sources using Flume and Sqoop.
- Monitor the data streaming between web sources and HDFS.
- Monitor the Hadoop cluster functioning through monitoring tools.
- Close monitoring and analysis of the MapReduce job executions on cluster at task level.
- Inputs to development regarding the efficient utilization of resources like memory and CPU utilization based on the running statistics of Map and Reduce tasks
- Changes to the configuration properties of the cluster based on volume of the data being processed and performance of the cluster.
- Handle the upgrades and Patch updates.
- Involved in loading data from UNIX file system to HDFS.
- Installed and configured Hive and also written Hive QL scripts.
- Responsible to manage data coming from different sources.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
- Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
- Extensive usage of Struts, HTML, CSS, JSP, JQuery, AJAX and JavaScript for interactive pages.
- Used Ganglia to monitor the cluster around the clock.
- Assist the team in their development & deployment activities.
- Set up automated processes to analyze the System and Hadoop log files for predefined errors and send alerts to appropriate groups.
- Set up and manage HA NameNode and NameNode federation using Apache 2.0 to avoid single point of failures in large clusters.
- Set up the checkpoints to gathering the system statistics for critical set ups.
- Discussions with other technical teams on regular basis regarding upgrades, Process changes, any Special processing and feedback.
Environment: Distcp,Kerberos,Ranger,HDFS,Pig,Hive,Sqoop,ZooKeeper,Oozie,Java,JEE,Spark,Solr,HBase.
Confidential
Java/J2EE developer
Responsibilities:
- Involved in designing Class and Sequence diagrams with UML and Data flow diagrams.
- Implemented MVC architecture using Struts framework to get the Free Quote.
- Designed and developed front end using JSP, Struts (tiles), XML, JavaScript, and HTML.
- Used Struts tag libraries to create JSP.
- Implemented Spring MVC, dependency Injection (DI) and aspect oriented programming (AOP) features along with Hibernate.
- Experienced with implementing navigation using Spring MVC.
- Used Hibernate for object-relational mapping persistence.
- Implemented message driven beans to get from queues to send again to support team using MSend commands.
- Experienced with hibernate core interfaces like configuration, session factory, transactional and criteria interfaces.
- Reviewed the requirements and Involved in database design for new requirements
- Wrote Complex SQL queries to perform various database operations using TOAD.
- Java Mail API was used to notify the Agents about the free quote and for sending Email to the Customer with Promotion Code for validation.
- Involved in testing using Junit.
- Performed application development using Eclipse and Web Sphere Application Server for deployment.
- Used SVN for version control.
Environment: Java, Spring, Hibernate, Jms, Web Services, Ejb, Sql/PlSql, Html, Css, Jsp, java script, Ant, Junit, Web sphere,
Confidential
Java Developer
Responsibilities:
- Implemented server side programs by using Servlets and JSP
- Designed, developed and validated User Interface using HTML, Java Script, XML and CSS
- Implemented MVC using Struts Framework
- Implemented Controller Servlet to handle the access to database
- Participated in code walkthroughs, Debugging and defect fixing
- Involved in the co-ordination of end to end production release process
- Used SVN for versioning control system
- Developed PL/SQL stored procedures and triggers
- Used JDBC prepared statements to call from Servlets for database access
- Designed and documented the stored procedures
- Extensively used HTML for web based design
- Involved in writing JUnit Test Cases and done unit testing for various components
- Worked on database interaction layer for insertions, updating and retrieval operations of data from oracle database by writing stored procedures
- Used Spring Framework for Dependency Injection and integrated with Hibernate
- Used Log4J for any errors in the application
Environment: Java, J2EE, JSP, Servlets, HTML, DHTML, XML, JavaScript, Struts, Eclipse, WebLogic, PL/SQL and Oracle
