We provide IT Staff Augmentation Services!

Sr. Hadoop Developer Resume

5.00/5 (Submit Your Rating)

Boston, MA

SUMMARY:

  • Highly skilled Java/Hadoop developer with strong background in programming and 5+ years of IT experience in developing software in Hadoop and Java based technologies.
  • Good understanding of complex processing needs of big data and has experience developing codes and modules to address those needs.
  • Hands - on coder with reliable and high quality code and will work well under pressure.
  • Solid understanding of Hadoop Distributed File System.
  • Good experience withMapReduce (MR), Hive, Pig, HBase, Sqoop, Oozie,Flume, Spark, Zookeeper for data extraction, processing, storage and analysis.
  • In-depth understanding on how MapReduce works and Hadoopinfrastructure.
  • Experience in developing custom MapReduce Programs in Java using ApacheHadoop for analyzing BigData as per the requirement. 
  • Extensively worked on Hive for ETL Transformations and optimized Hive Queries
  • Worked with relational database systems (RDBMS) such as MySQL, MSSQL, Oracle and NoSQL database systems like HBase and Cassandra.
  • Good Knowledge on HadoopCluster architecture and monitoring the cluster .
  • Used Shell Scripting to move log files into HDFS
  • Strong Hands on experience in MVC frameworks, Struts and Spring MVC.
  • Wrote complex Hive, Pig UDF’s using python scripting.
  • Good in Designing and developing the Data Access Layer modules with the help of Hibernate Framework for the new functionalities.
  • Extensively experience in working on IDEs like Eclipse, Net Beans and Edit Plus.
  • Working knowledge of Agile and waterfall development models.
  • Extensively used Java and J2EE technologies like Core Java, Java Beans, Servlet, JSP, Struts, spring, Hibernate, JDBC, JSONObject, and Design Patterns.
  • Experienced in Application Development using Java, J2EE, JSP, Servlets, Struts, RDBMS, Tag Libraries, JDBC, Hibernate, XML and Linux shell scripting.
  • Worked with different software version control, bug tracking and code review systems like CVS, Source Depot, Bugger, TFS and Code Flow.

TECHNICAL SKILLS:

Big Data: Hadoop, Map Reduce, HDFS, Hive, HBase, Pig, Sqoop, Oozie, Zookeeper, Flume, YARN, Storm, Spark, Mongo DB, Kafka, Cassandra and Impala.

Java/J2EE: Java, J2EE, JSP, JavaScript, Servlets, JDBC, Java Beans, JMS, EJB.

Frameworks: Apache Struts, Hibernate, Spring, MVC.

Languages: Core Java, J2EE, C, SQL, PL/SQL, UML, Scala.

WebService/ Technologies: REST, SOAP, JSP, JavaScript, XML, HTML, CSS, AJAX and JSON.

Databases: NoSQL, SQL Server, MySQL, Oracle, PL/SQL.

Web/Application Servers: Apache Tomcat, Web Logic, Web Sphere, JBOSS.

Tools: Eclipse, NetBeans and Edit Plus.

Build Tools: UML, Design Patterns, Maven, Ant.

Operating System: Linux,Windows XP/7/8.

Hadoop Distributions: Cloudera, Horton works, MapR.

Version Controls: SVN, TFS

PROFESSIONAL EXPERIENCE:

Confidential

Boston, MA

Sr. Hadoop Developer

Responsibilities:

  • Extracted large datasets from Teradata and Oracle servers into HDFS and vice versa using Sqoop.
  • Performed various optimization techniques likes partioning and bucketing on Hive tables.
  • Analyzed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most visited page on website.
  • Written Hive UDF, UDTF and UDAFs in java to extend Hive functionality depending on the requirements.
  • Responsible for automating number of Sqoop, Hive and Pig scripts using Oozie Workflow Scheduler.
  • Used Impala to pull Hive table data for faster query processing and pushed the results to Cassandra.
  • Experienced with working on Avro and Parquet Data files using Avro and Parquet Serialization.
  • Experienced in developing Spark applications using Scala on different data formats like Text file, CSV file.
  • Developed Spark applications using Scala by importing Spark-SQL API's like DataFrames and DataSets API's.
  • Extensively worked on developing POC using Spark streaming application connecting to Kafka's consumer and pulling the data into Spark core for analysis.
  • Configured Spark Streaming to receive real time data from the Kafka and store the stream data to HDFS and performed analysis using Spark Core API and Scala programming.
  • Analysis and development of Spark Cassandra connector to load data from flat file to Cassandra.

Environment: Hadoop, Yarn, Sqoop, Flume, Hive, Pig, Spark, Oozie, Impala, Cassandra, Kafka, Scala, Core Java.

Confidential

Somerset, NJ

Hadoop Developer

Responsibilities:

  • Setup scripts to fetch data from various ftp server locations and copy them into HDFS folder corresponding to the client.
  • Defined client-agnostic formats for different kinds of data we receive from the clients.
  • Wrote Pig UDFs to pre-process the data received from various clients, and transform them to the required formats.
  • Specified numerous Pig relations to map various fields in the data set.
  • Developed various Pig Latin scripts to join, groupdifferent kinds of data to construct relevant records according to the functional requirement.
  • Developed MapReduce programs for analyzing the data, in cases where Pig scripts performance is not satisfactory.
  • Used Flume to process real time processing data.
  • Utilized HCATALOG to access Hive tables metadata from Pig scripts and MapReduce jobs.
  • Used Hive to do analysis on the data and identify different correlations.
  • Written AdhocHiveQL queries to process data and generate reports.
  • Involved in HDFS maintenance and administering it through Hadoop-Java API.
  • Worked on importing and exporting data from Oracle and DB2 into HDFS and HIVE using Sqoop.
  • Configured MySQL Database to store Hive metadata.
  • Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
  • Written Hive queries for data analysis to meet the business requirements.
  • Automated all the jobs, for pulling data from FTP server to load data into Hive tables, using Oozie workflows.
  • Collecting and aggregating large amounts of log data using ApacheFlume and staging data in HDFS for further analysis.
  • Involved in creating Hive tables and working on them using Hive QL.
  • Involved in maintaining and monitoring clusters.
  • Extracted the data from MySQL, Oracle, Sql Server using Sqoop and loaded data into Cassandra.
  • Used Cassandra nodetool to manage Cassandra cluster.
  • Designing and implementing semi-structured data analytics platform leveraging Hadoop, with Solr.
  • Implemented test scripts to support test driven development and continuous integration.
  • Automated the jobs to pull the data from ftp servers to HDFS using Oozie workflows and enabled email alerts for communication in case of any failure.
  • Performed unit testing of MapReduce jobs using MRUnit.
  • Worked closely with the Data Analyst to identify the business aspects for analysis.
  • Took part in managing and reviewing log files.
  • Involved in set up of Oracle R connector for Hadoop so that data analyst can use data in HDFS to perform analytics.
  • Actively took part in scrum meetings to discuss the progress of the deliverables.

Environment: CDH4, HDFS, Cloudera Manager, MapReduce, Linux, Putty, Pig, Hive, Oozie, MRUnit, Shell scripting, Eclipse Luna, Java, VersionOne.

Confidential 

Nashville, TN

Java Developer

Responsibilities:

  • Involved in design and implementation of server side programming.
  • Coordinating with Project Manager for getting the requirements and developing the code to support new applications.
  • Involved in gathering requirements, analyzed them and prepared high level documents.
  • Participated in all client meetings to understand the requirements.
  • Actively involved in designing and data modelling using Rational Rose Tool (UML).
  • Involved in the design of the SPACE database.
  • Designed and development of User Interfaces, Menus using HTML, JSP, JSP Custom Tag, Java Script.
  • Implemented User Interface Using Spring Tiles framework.
  • Involved in integrating system with BT’s systems like GTC, CSS, OR through e Link Hub, and IBM MQ series.
  • Developed, Deployed and tested JSPs, Servlets in WebLogic.
  • Used Eclipse as IDE tool and integrated Web Logic with Eclipse to develop & deploy the applications.

Environment: Core Java/J2EE, Servlet, JSP, Struts (MVC2), HTML, XML, CSS, Ajax, Linux, MVC, MySQL, .Net Beans, My-Eclipse, Apache Tomcat.

Confidential

Knoxville

Java Developer

Responsibilities:

  • Involved in gathering business requirements, analyzing the project and created UML diagrams such as Use Cases, Class Diagrams, Sequence Diagrams and flowcharts for the optimization Module using Microsoft Visio.
  • Designed and developed Optimization UI screens for Rate Structure, Operating Cost, Temperature and Predicted loads using JSF myfaces, JSP, JavaScript and HTML.
  • Configured faces-config.xml for the page navigation rules and created managed and backing beans for the Optimization module.
  • Developed JSP web pages for rate Structure and Operating cost using JSF HTML and JSF CORE tags library.
  • Designed and developed the framework for the IMAT application implementing all the six phases of JSF life cycle and wrote Ant build, deployment scripts to package and deploy on JBoss application server.
  • Designed and developed Simulated annealing algorithm to generate random Optimization schedules and developed neural networks for the CHP system using Session Beans.
  • Integrated EJB 3.0 with JSF and managed application state management, business process management (BPM) using JBoss Seam.
  • Wrote Angular.JS controllers, views, and services for new website features.
  • Developed Cost function to calculate the total cost for each CHP Optimization schedule generated by the Simulated Annealing algorithm using EJBs.
  • Implemented spring web flow for the Diagnostics Module to define page flows with actions and views and created POJOs and used annotations to map them to SQL Server database using EJB.
  • Wrote DAO classes, EJB 3.0 QL queries for Optimization schedule and CHP data retrievals from SQL Server database.
  • Used Eclipse as IDE tool to develop the application and JIRA for bug and issue tracking
  • Created combined deployment descriptors using XML for all the session and entity beans.
  • Wrote JSF and JavaScript validations to validate data on the UI for Optimization and Diagnostics and Developed Web Services to have access to the external system (WCC) for the
  • Designed and coded application components in an agile environment utilizing a test driven development approach.
  • Skilled in test driven development and agile development.
  • Created technical design document for the Diagnostics Module and Optimization module covering Cost function and Simulated Annealing approach.
  • Involved in code reviews and performed version guidelines.

Environment: Java 1.5, J2EE, Microsoft Vision, EJB 3.0, JSP, JSF, JBoss Seam, JIRA, Web Services, JMS, JavaScript, Angular.js, HTML, ANT, Agile, JUnit, JBoss 4.2.2, MS SQL Server 2005, My ECLIPSE 6.0.1.

We'd love your feedback!