Big Data / Hadoop Developer Resume
Scottsdale, AZ
SUMMARY:
- Over 6 years of professional IT experience and certified Hadoop developer with over 4 years of Hadoop ecosystem experience. Efficient in Data design and development using ETL methodologies.
- In depth and extensive knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN, Resource Manager, Node Manager and Map Reduce.
- Hands on experience in using Hadoop ecosystem components Hive, Pig, Oozie, Sqoop, Flume, HUE, ZooKeeper.
- Extensive experience in processing data using Spark (Scala, SparkSQL) and Spark Streaming.
- Hands on experience in processing real time data streams using Spark Streaming.
- Integrated Spark Streaming with Kafka, Amazon Kinesis and Cassandra.
- Familiarity and Good Understanding in Kafka and Storm.
- Experience in processing vast data sets by performing structural modifications, to cleanse both structured and semi - structured data using MapReduce programs in JAVA, HiveQL and Pig Latin.
- Experience in performing specialized JOINS with PIG, HIVE and Map-Reduce java.
- Worked with multiple file Input Formats such as TextFile, KeyValue and SequenceFileinput format.
- Experience in working with multiple file formats JSON, XML, Sequence Files and RC Files.
- Hands on experience in writing custom Key, Value, InputFormat, RecordReader, Combiner and Partitioner in Java.
- Expertise in optimization of mapreduce algorithms using Combiners, Partitioners and Distributed Cache to deliver best results.
- Solid Experience in optimizing the Hive queries using Partitioning and Bucketing techniques, which controls the data distribution, to enhance performance.
- Experience in writing custom UDF’s for PIG & HIVE, extending their library as required for working on data in non-standard formats.
- Extensive experience in using Hive UDAF’s, UDTF’s and Hive SerDes.
- Hands on experience in developing Pig Scripts to perform data transformation operations, by implementing various functions, for loading and evaluating data in the relations.
- Experience in Importing and Exporting data from different databases like MySql, Oracle into HDFS and Hive using Sqoop.
- Loading log data into HDFS by collecting and aggregating the data from various sources using Flume.
- Familiar with writing Oozie workflows and Job Controllers for job automation.
- Configured the Hadoop cluster in Local (Standalone), Pseudo-Distributed, Fully-Distributed mode with the use of Apache, Cloudera distributions and Cloudera manager.
- Experience in Setting up NameNode High Availability (Quorum based) with Automatic Failover, to eliminate single point of failure for Hadoop clusters.
- Experience in developing Shell Scripts for system management and for automating routine tasks.
- Good exposure to Namenode Federation and MapReduce 2.0 (MRv2) or YARN.
- Experience in integrating Hadoop with Enterprise BI and DW tools Tableau and Pentaho. Familiarity with NoSQL databases HBase and Cassandra.
- Well versed knowledge and experience in public cloud environment - Amazon Web Services (AWS), RackSpace and on private cloud infra structure - OpenStack cloud platform.
- Object Oriented Programming using Java and J2EE related technologies.
- Experience with front end technologies HTML (5), CSS, JavaScript, XML and JQuery. Worked extensively on different databases Oracle, MySQL and have good database programming experience with SQL.
- Experience in writing java programs to parse JSON files and XML files using SAX and DOM parsers.
- Has in depth knowledge in Digital Marketing and Travel domains.
TECHNICAL SKILLS:
Programming Languages: Java, SQL, UNIX Shell Scripting, Scala, Python
Web Technologies: HTML(5), CSS, JavaScript, XML, JQuery, Ajax, PHP
BIG Data: Apache Hadoop, Hive, Pig, Java Map Reduce, Spark, Spark StreamingImpala, Sqoop, Flume, Kafka, Oozie, Hue, ZooKeeper, HBaseCassandra
Databases: Oracle, MySQL, MS SQL Server, MS-Access
Java Technologies & Frameworks: JDBC, JSP, Servlets, Spring, Hibernate
Web Servers: Tomcat, WebSphere, WebLogic, WAMP, LAMP
Tools: & Technologies: Maven, Ant, Eclipse, JDeveloper, VMware vsphere, MS Visual StudioPuTTY, WinSCP
Reporting & ETL Tools: Tableau, Pentaho
Cloud Platforms: Amazon Web Services, Open Stack, RackSpace
Monitoring&Administration tools: Cloudera Manager, Nagios, Ganglia
Version Control: GitHub, SVN, Clear Case
Operating Systems: Unix, Linux(RedHat, CentOs, Ubuntu), Windows, Mac
Domain knowledge: Travel Data, Digital Marketing and Insurance
PROFESSIONAL EXPERIENCE:
Confidential, Scottsdale, AZ
Big Data / Hadoop Developer
Responsibilities:
- Highly Experienced in batch-processing of data using PIG and Hive.
- Used various built-in functions and regex operations to extract the sensible data also developed Pig UDF’s in java used to derive business metrics.
- Optimized the processing by using specialized joins in Pig and bucketing in Hive.
- Worked on generating reports using SPARK SQL for fast processing of data.
- Highly Worked on unix shell scripting to handle day to day jobs.
- Used Oozie workflows to automate various jobs i.e., Pig, Hql and shell scripts.
- Also used NiFi flow to automate the process of transmitting data through MFT.
- Worked on developing different kinds of reports using different loaders and delimiters. Involved in end to end development and deployment of reports for various clients.
- Worked on development of Kafka-Spark streaming integration to process real-time data.
- Developed Spark streaming pipeline in java to parse JSON data and to store in Hive tables.
- Involved and contributed in architectural discussions and handled offshore team.
- Worked closely with Data warehouse team and business analysts to reduce the gaps between the requirement gathering.
- Used JIRA tracker for daily work logs and SDLC management.
Environment: Hadoop Architecture, HDFS, MapReduce, Hive, Pig, Spark, Kafka, Sqoop, Nifi, Oozie, ctrl-M, Java, Burst,MySQL, CentOS, Unix, PuTTY, JIRA.
Confidential
Big Data / Hadoop Developer
Responsibilities:
- Extensive experience in processing data using Spark (Scala, SparkSQL) and Spark Streaming.
- Hands on experience in processing real time data streams using Spark Streaming.
- Integrated Spark Streaming with Kafka, Flume, Amazon Kinesis and Cassandra.
- Familiarity and Good Understanding in Kafka and Storm.
- Worked on developing python scripts using REST and SOAP APIs.
- Extensive experience in handling transient clusters and persistent clusters using shell script on AWS.
- Developed multiple hive scripts and redshift scripts for several workflows. Used Sqoop to efficiently transfer data between databases and HDFS.
- Involved in architecting Hadoop clusters Translation of functional and technical requirements into detailed architecture and design.
- Connected qlikview reporting tool to redshift and generated reports.
- Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
- Developed Unix Shell scripts to automate the cluster installation.
- Involved in developing reports in Tableau and apache zeppelin
Environment: Hadoop Architecture, HDFS, MapReduce, Hive, Pig, Spark, Kafka, Storm, Nifi, Sqoop, Oozie, Flume, Java, MySQL, Tableau, CentOS, Windows, Unix, AWS, PuTTY.
Confidential, DC
Big Data / Hadoop Developer
Responsibilities:
- Designed and implemented Customization of Keys, Values, Partitioners, Combiners, InputFormats and RecordReaders in JAVA.
- Implemented UDFS, UDAFS, UDTFS in java for hive to process the data that can’t be performed using Hive inbuilt functions.
- Worked with multiple Input Formats such as TextFile, KeyValue, SequenceFile and NLine input format.
- Worked with multiple file formats JSON, XML, Sequence Files and RC Files.
- Deployed and configured Flume agents to stream log events into HDFS for analysis. Transformed the log files into structured data using Hive SerDe’s and Pig Loaders.
- Parsed JSON and XML files in PIG using Pig Loader functions and extracted meaningful information from Pig Relations by providing a regex using the built-in functions in Pig.
- Involved in creating Hive Internal and External tables, loading data and writing hive queries which will run internally in map reduce way.
- Analyzed JSON and XML files using Hive Built in functions and SerDe’s.
- Optimizing the Hive queries using Partitioning and Bucketing techniques, for controlling the data distribution.
- Used Sqoop to efficiently transfer data between databases and HDFS.
- Implemented complex map reduce programs to perform joins on the Map side using Distributed Cache in Java.
- Worked on complex data types Array, Map and Struct in Hive.
- Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
- Developing Scripts and Batch Jobs to schedule various Hadoop Programs. Familiarity in using NoSQL database, HBase on top of HDFS.
- Worked with BI teams in generating the reports in Tableau and designing ETL workflows on Pentaho.
Environment: Hadoop, HDFS, MapReduce, Java, Hive, Pig, Sqoop, Flume, Oozie, Hue, Cloudera Manager, HBase, Tableau, Pentaho, CentOS, Windows, PuTTY, MySQL.
Confidential
Big Data / Hadoop Developer
Responsibilities:
- Developed Map-Reduce programs in JAVA, HIVE and PIG to validate and cleanse the data in HDFS, obtained from heterogeneous data sources, to make it suitable for analysis.
- Used a 20 node cluster with Cloudera Hadoop distribution on Amazon EC2 for backup. Worked on streaming the data into HDFS from web servers using Flume.
- Designed and implemented Hive and Pig UDF’s for evaluation, filtering, loading and storing of data.
- Experience in writing Nested foreach in PigLatin, for implementing complex business logic in the process of data mining.
- Performed Fragment-Replicate Joins (Map side joins) in Pig, which implements Distributed Cache.
- The Hive tables created as per requirement were Internal or External tables defined with appropriate Static and Dynamic partitions, intended for efficiency.
- Wrote Pig Scripts to generate MapReduce jobs and performed ETL procedures on the data in HDFS.
- Used several Built in functions and PiggyBank UDF’s for data transformations in Pig. Implemented Lateral View in conjunction with UDTFs in Hive.
- Worked extensively with Sqoop for importing and exporting data from MySQL into HDFS and Hive.
- Performed complex Joins on the tables in Hive.
- Load and transform large sets of structured, semi structured using Hive and Impala.
- Connected Hive and Impala to Tableau reporting tool and generated graphical reports.
- Understanding the existing Enterprise data warehouse set up and provided design and architecture suggestion converting to Hadoop using MR, HIVE, SQOOP and Pig Latin.
Environment: Hadoop(CDH), HDFS, MapReduce, Hive, Pig, Impala, Sqoop, Oozie, Flume, Hue, Java, MySQL, Tableau, CentOS, Unix, AWS, PuTTY.
Confidential
Sr. Java/J2EE Developer
Responsibilities:
- Actively involved in analyzing and collecting user requirements
- Participated in Server side and Client-side programming
- Developed and tested the DAO layer for CRUD operation.
- Involved in complete Software Development Life Cycle (SDLC) with Object Oriented approach of client’s business process and continuous client feedback.
- Used HTML, CSS, JavaScript, JQuery, Ajax for Front End Development.
- Implemented the web-based application using Spring framework.
- Involved in coding of JSP pages, for the presentation of data, on the View layer in MVC Architecture.
- Used Spring dependency injection and Spring-Hibernate Integration.
- Responsible for writing JavaScript for validation in client side.
- Parsed through JSON and XML files using JSONParser, SAX and DOM Parsers.
- Wrote database queries using SQL for accessing, manipulating and updating Oracle Database.
Environment: Java/J2EE, Web Services, SVN, Eclipse, Spring, Hibernate, JSON, XML, JSP, JavaScript, HTML, CSS, JQuery, AJAX, Tomcat Server, Oracle 10g, Windows.
Confidential
Jr. Java/J2EE Developer
Responsibilities:
- Worked with a team of 4 members to deliver the end to end web application for an educational institution.
- Involved in analyzing, designing, coding and testing.
- Designed and implemented UI layer using HTML, JavaScript and JSP Servlets.
- Connected java web applications to MySQL by using JDBC connectors in Eclipse IDE.
- Developed the application using Spring MVC framework.
- Involved in the migration of independent parts of the system to use Hibernate for the implementation of DAO.
- Involved in configuring and deploying the application on Tomcat Server.
- Involved on the back end to modify business logic by making enhancements.
Environment: Java, J2EE, Spring, Hibernate, Eclipse, MVC, JUnit, JSP, DHTML, JavaScript, Ajax, Web Services, Tomcat, Rational Rose, SOAP, Windows, UNIX.
