We provide IT Staff Augmentation Services!

Big Data / Hadoop Developer Resume

4.00/5 (Submit Your Rating)

Scottsdale, AZ

SUMMARY:

  • Over 6 years of professional IT experience and certified Hadoop developer with over 4 years of Hadoop ecosystem experience. Efficient in Data design and development using ETL methodologies.
  • In depth and extensive knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN, Resource Manager, Node Manager and Map Reduce.
  • Hands on experience in using Hadoop ecosystem components Hive, Pig, Oozie, Sqoop, Flume, HUE, ZooKeeper.
  • Extensive experience in processing data using Spark (Scala, SparkSQL) and Spark Streaming.
  • Hands on experience in processing real time data streams using Spark Streaming.
  • Integrated Spark Streaming with Kafka, Amazon Kinesis and Cassandra.
  • Familiarity and Good Understanding in Kafka and Storm.
  • Experience in processing vast data sets by performing structural modifications, to cleanse both structured and semi - structured data using MapReduce programs in JAVA, HiveQL and Pig Latin.
  • Experience in performing specialized JOINS with PIG, HIVE and Map-Reduce java.
  • Worked with multiple file Input Formats such as TextFile, KeyValue and SequenceFileinput format.
  • Experience in working with multiple file formats JSON, XML, Sequence Files and RC Files.
  • Hands on experience in writing custom Key, Value, InputFormat, RecordReader, Combiner and Partitioner in Java.
  • Expertise in optimization of mapreduce algorithms using Combiners, Partitioners and Distributed Cache to deliver best results.
  • Solid Experience in optimizing the Hive queries using Partitioning and Bucketing techniques, which controls the data distribution, to enhance performance.
  • Experience in writing custom UDF’s for PIG & HIVE, extending their library as required for working on data in non-standard formats.
  • Extensive experience in using Hive UDAF’s, UDTF’s and Hive SerDes.
  • Hands on experience in developing Pig Scripts to perform data transformation operations, by implementing various functions, for loading and evaluating data in the relations.
  • Experience in Importing and Exporting data from different databases like MySql, Oracle into HDFS and Hive using Sqoop.
  • Loading log data into HDFS by collecting and aggregating the data from various sources using Flume.
  • Familiar with writing Oozie workflows and Job Controllers for job automation.
  • Configured the Hadoop cluster in Local (Standalone), Pseudo-Distributed, Fully-Distributed mode with the use of Apache, Cloudera distributions and Cloudera manager.
  • Experience in Setting up NameNode High Availability (Quorum based) with Automatic Failover, to eliminate single point of failure for Hadoop clusters.
  • Experience in developing Shell Scripts for system management and for automating routine tasks.
  • Good exposure to Namenode Federation and MapReduce 2.0 (MRv2) or YARN.
  • Experience in integrating Hadoop with Enterprise BI and DW tools Tableau and Pentaho. Familiarity with NoSQL databases HBase and Cassandra.
  • Well versed knowledge and experience in public cloud environment - Amazon Web Services (AWS), RackSpace and on private cloud infra structure - OpenStack cloud platform.
  • Object Oriented Programming using Java and J2EE related technologies.
  • Experience with front end technologies HTML (5), CSS, JavaScript, XML and JQuery. Worked extensively on different databases Oracle, MySQL and have good database programming experience with SQL.
  • Experience in writing java programs to parse JSON files and XML files using SAX and DOM parsers.
  • Has in depth knowledge in Digital Marketing and Travel domains.

TECHNICAL SKILLS:

Programming Languages: Java, SQL, UNIX Shell Scripting, Scala, Python

Web Technologies: HTML(5), CSS, JavaScript, XML, JQuery, Ajax, PHP

BIG Data: Apache Hadoop, Hive, Pig, Java Map Reduce, Spark, Spark StreamingImpala, Sqoop, Flume, Kafka, Oozie, Hue, ZooKeeper, HBaseCassandra

Databases: Oracle, MySQL, MS SQL Server, MS-Access

Java Technologies & Frameworks: JDBC, JSP, Servlets, Spring, Hibernate

Web Servers: Tomcat, WebSphere, WebLogic, WAMP, LAMP

Tools: & Technologies: Maven, Ant, Eclipse, JDeveloper, VMware vsphere, MS Visual StudioPuTTY, WinSCP

Reporting & ETL Tools: Tableau, Pentaho

Cloud Platforms: Amazon Web Services, Open Stack, RackSpace

Monitoring&Administration tools: Cloudera Manager, Nagios, Ganglia

Version Control: GitHub, SVN, Clear Case

Operating Systems: Unix, Linux(RedHat, CentOs, Ubuntu), Windows, Mac

Domain knowledge: Travel Data, Digital Marketing and Insurance

PROFESSIONAL EXPERIENCE:

Confidential, Scottsdale, AZ

Big Data / Hadoop Developer

Responsibilities:

  • Highly Experienced in batch-processing of data using PIG and Hive.
  • Used various built-in functions and regex operations to extract the sensible data also developed Pig UDF’s in java used to derive business metrics.
  • Optimized the processing by using specialized joins in Pig and bucketing in Hive.
  • Worked on generating reports using SPARK SQL for fast processing of data.
  • Highly Worked on unix shell scripting to handle day to day jobs.
  • Used Oozie workflows to automate various jobs i.e., Pig, Hql and shell scripts.
  • Also used NiFi flow to automate the process of transmitting data through MFT.
  • Worked on developing different kinds of reports using different loaders and delimiters. Involved in end to end development and deployment of reports for various clients.
  • Worked on development of Kafka-Spark streaming integration to process real-time data.
  • Developed Spark streaming pipeline in java to parse JSON data and to store in Hive tables.
  • Involved and contributed in architectural discussions and handled offshore team.
  • Worked closely with Data warehouse team and business analysts to reduce the gaps between the requirement gathering.
  • Used JIRA tracker for daily work logs and SDLC management.

Environment: Hadoop Architecture, HDFS, MapReduce, Hive, Pig, Spark, Kafka, Sqoop, Nifi, Oozie, ctrl-M, Java, Burst,MySQL, CentOS, Unix, PuTTY, JIRA.

Confidential

Big Data / Hadoop Developer

Responsibilities:

  • Extensive experience in processing data using Spark (Scala, SparkSQL) and Spark Streaming.
  • Hands on experience in processing real time data streams using Spark Streaming.
  • Integrated Spark Streaming with Kafka, Flume, Amazon Kinesis and Cassandra.
  • Familiarity and Good Understanding in Kafka and Storm.
  • Worked on developing python scripts using REST and SOAP APIs.
  • Extensive experience in handling transient clusters and persistent clusters using shell script on AWS.
  • Developed multiple hive scripts and redshift scripts for several workflows. Used Sqoop to efficiently transfer data between databases and HDFS.
  • Involved in architecting Hadoop clusters Translation of functional and technical requirements into detailed architecture and design.
  • Connected qlikview reporting tool to redshift and generated reports.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
  • Developed Unix Shell scripts to automate the cluster installation.
  • Involved in developing reports in Tableau and apache zeppelin

Environment: Hadoop Architecture, HDFS, MapReduce, Hive, Pig, Spark, Kafka, Storm, Nifi, Sqoop, Oozie, Flume, Java, MySQL, Tableau, CentOS, Windows, Unix, AWS, PuTTY.

Confidential, DC

Big Data / Hadoop Developer

Responsibilities:

  • Designed and implemented Customization of Keys, Values, Partitioners, Combiners, InputFormats and RecordReaders in JAVA.
  • Implemented UDFS, UDAFS, UDTFS in java for hive to process the data that can’t be performed using Hive inbuilt functions.
  • Worked with multiple Input Formats such as TextFile, KeyValue, SequenceFile and NLine input format.
  • Worked with multiple file formats JSON, XML, Sequence Files and RC Files.
  • Deployed and configured Flume agents to stream log events into HDFS for analysis. Transformed the log files into structured data using Hive SerDe’s and Pig Loaders.
  • Parsed JSON and XML files in PIG using Pig Loader functions and extracted meaningful information from Pig Relations by providing a regex using the built-in functions in Pig.
  • Involved in creating Hive Internal and External tables, loading data and writing hive queries which will run internally in map reduce way.
  • Analyzed JSON and XML files using Hive Built in functions and SerDe’s.
  • Optimizing the Hive queries using Partitioning and Bucketing techniques, for controlling the data distribution.
  • Used Sqoop to efficiently transfer data between databases and HDFS.
  • Implemented complex map reduce programs to perform joins on the Map side using Distributed Cache in Java.
  • Worked on complex data types Array, Map and Struct in Hive.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
  • Developing Scripts and Batch Jobs to schedule various Hadoop Programs. Familiarity in using NoSQL database, HBase on top of HDFS.
  • Worked with BI teams in generating the reports in Tableau and designing ETL workflows on Pentaho.

Environment: Hadoop, HDFS, MapReduce, Java, Hive, Pig, Sqoop, Flume, Oozie, Hue, Cloudera Manager, HBase, Tableau, Pentaho, CentOS, Windows, PuTTY, MySQL.

Confidential

Big Data / Hadoop Developer

Responsibilities:

  • Developed Map-Reduce programs in JAVA, HIVE and PIG to validate and cleanse the data in HDFS, obtained from heterogeneous data sources, to make it suitable for analysis.
  • Used a 20 node cluster with Cloudera Hadoop distribution on Amazon EC2 for backup. Worked on streaming the data into HDFS from web servers using Flume.
  • Designed and implemented Hive and Pig UDF’s for evaluation, filtering, loading and storing of data.
  • Experience in writing Nested foreach in PigLatin, for implementing complex business logic in the process of data mining.
  • Performed Fragment-Replicate Joins (Map side joins) in Pig, which implements Distributed Cache.
  • The Hive tables created as per requirement were Internal or External tables defined with appropriate Static and Dynamic partitions, intended for efficiency.
  • Wrote Pig Scripts to generate MapReduce jobs and performed ETL procedures on the data in HDFS.
  • Used several Built in functions and PiggyBank UDF’s for data transformations in Pig. Implemented Lateral View in conjunction with UDTFs in Hive.
  • Worked extensively with Sqoop for importing and exporting data from MySQL into HDFS and Hive.
  • Performed complex Joins on the tables in Hive.
  • Load and transform large sets of structured, semi structured using Hive and Impala.
  • Connected Hive and Impala to Tableau reporting tool and generated graphical reports.
  • Understanding the existing Enterprise data warehouse set up and provided design and architecture suggestion converting to Hadoop using MR, HIVE, SQOOP and Pig Latin.

Environment: Hadoop(CDH), HDFS, MapReduce, Hive, Pig, Impala, Sqoop, Oozie, Flume, Hue, Java, MySQL, Tableau, CentOS, Unix, AWS, PuTTY.

Confidential

Sr. Java/J2EE Developer

Responsibilities:

  • Actively involved in analyzing and collecting user requirements
  • Participated in Server side and Client-side programming
  • Developed and tested the DAO layer for CRUD operation.
  • Involved in complete Software Development Life Cycle (SDLC) with Object Oriented approach of client’s business process and continuous client feedback.
  • Used HTML, CSS, JavaScript, JQuery, Ajax for Front End Development.
  • Implemented the web-based application using Spring framework.
  • Involved in coding of JSP pages, for the presentation of data, on the View layer in MVC Architecture.
  • Used Spring dependency injection and Spring-Hibernate Integration.
  • Responsible for writing JavaScript for validation in client side.
  • Parsed through JSON and XML files using JSONParser, SAX and DOM Parsers.
  • Wrote database queries using SQL for accessing, manipulating and updating Oracle Database.

Environment: Java/J2EE, Web Services, SVN, Eclipse, Spring, Hibernate, JSON, XML, JSP, JavaScript, HTML, CSS, JQuery, AJAX, Tomcat Server, Oracle 10g, Windows.

Confidential

Jr. Java/J2EE Developer

Responsibilities:

  • Worked with a team of 4 members to deliver the end to end web application for an educational institution.
  • Involved in analyzing, designing, coding and testing.
  • Designed and implemented UI layer using HTML, JavaScript and JSP Servlets.
  • Connected java web applications to MySQL by using JDBC connectors in Eclipse IDE.
  • Developed the application using Spring MVC framework.
  • Involved in the migration of independent parts of the system to use Hibernate for the implementation of DAO.
  • Involved in configuring and deploying the application on Tomcat Server.
  • Involved on the back end to modify business logic by making enhancements.

Environment: Java, J2EE, Spring, Hibernate, Eclipse, MVC, JUnit, JSP, DHTML, JavaScript, Ajax, Web Services, Tomcat, Rational Rose, SOAP, Windows, UNIX.

We'd love your feedback!