We provide IT Staff Augmentation Services!

Hadoop Developer Resume

4.00/5 (Submit Your Rating)

Woonsocket, RI

PROFESSIONAL SUMMARY:

  • Over 9+ years of experience in IT industry with 4 years of experience in Big Data implementing complete Hadoop solutions along with Scala and Java/J2EE.
  • Good working experience in using Apache Hadoop eco system components like MapReduce, HDFS, Hive, Sqoop, Pig, Oozie, Flume, HBase and Zoo Keeper.
  • Extensive experience working on spark in performing ETL using Spark Core, Spark - SQL and Real-time data processing using Spark Streaming.
  • Expertise in working with RDD’s, DataFrames and DataSet API’s.
  • Extensively worked with Kafka as middleware for real-time data pipelines.
  • Writing UDFs and integrating with Hive and Pig using Java.
  • Experience with Sequence files, AVRO and ORC file formats and compression.
  • Experience in Hadoop Distributions: Cloudera and Hortonworks,
  • Performed importing and exporting data into HDFS and Hive using Sqoop.
  • Experience in job workflow scheduling tools like Oozie.
  • Extensive knowledge in using SQL Queries for backend database analysis.
  • Strong knowledge in NOSQL column oriented databases like HBase, Cassandra and its integration with Hadoop cluster.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS) and vice-versa.
  • Led many Data Analysis & Integration efforts involving HADOOP along with ETL.
  • Hands on experience on Enterprise Data Lake to provide support for various uses cases including Analytics, processing, storing and Reporting of voluminous, rapidly changing, structured and unstructured data.
  • Extensive experience with SQL, PL/SQL and database concepts.
  • Transferred bulk data from RDBMS systems like Teradata into HDFS using Sqoop.
  • Experience in installation, configuration, supporting and managing Hadoop Clusters using Horton works, and Cloudera (CDH3, CDH4) distributions on Amazon web services (AWS)
  • Experience in analyzing data using Hive QL, Pig Latin, and custom MapReduce programs in Java.
  • Strong experience in analyzing large amounts of data sets writing PySpark scripts and Hive queries.
  • Worked in developing a Nifi flow prototype for data ingestion in HDFS.
  • Well-versed in Agile, other SDLC methodologies and can coordinate with owners and SMEs.
  • Worked on different operating systems like UNIX, Linux, and Windows
  • Diverse experience utilizing Java tools in business, Web, and client-server environments including Java Platform, Enterprise Edition (Java EE), Enterprise Java Bean (EJB), JavaServer Pages (JSP), Java Servlets (including JNDI), Struts, and Java database Connectivity (JDBC) technologies.
  • Fluid understanding of multiple programming languages, including C#, C, C++, JavaScript, HTML, and XML.
  • Experience in web application design using open source MVC , Spring and Struts Frameworks.

TECHNICAL SKILLS:

Hadoop Core Services: HDFS, Map Reduce, Spark, YARN

Hadoop Distribution: Horton works, Cloudera, Apache

NO SQL Databases: HBase, Cassandra

Hadoop Data Services: Hive, Pig, Sqoop, Flume

Hadoop Operational Services: Zookeeper, Oozie

Monitoring Tools: Cloudera Manager

Cloud Computing Tools: Amazon AWS

Languages: C, Java/J2EE, Python, SQL, PL/SQL, Pig Latin, HiveQL, Unix Shell Scripting

Java & J2EE Technologies: Core Java, Servlets, Hibernate, Spring, Struts

Application Servers: Web Logic, Web Sphere, JBoss, Tomcat

Databases: Oracle, MySQL, Postgress, Teradata

Operating Systems: UNIX, Windows, LINUX

Build Tools: Jenkins, Maven, ANT

Development Tools: Microsoft SQL Studio, Toad, Eclipse, NetBeans

Development methodologies: Agile/Scrum

Visualization and analytics too: Tableau Software, Qlik View

PROFESSIONAL EXPERIENCE:

Confidential - Woonsocket, RI

Hadoop Developer

Responsibilities:

  • Imported the data from various formats like JSON, Sequential, Text, CSV, AVRO and Parquet to HDFS cluster with compressed for optimization.
  • Worked on ingesting data from RDBMS sources like - Oracle, SQL Server and Teradata into HDFS using Sqoop.
  • Loaded all data-sets into Hive from Source CSV files using spark and Cassandra from Source CSV files using Spark
  • Built a Ingestion Framework that would ingest the files from SFTP to HDFS using Apache NIFI and ingest Financial data into HDFS
  • Created environment to access Loaded Data via spark SQL, through JDBC&ODBC (via Spark Thrift Server). Developed real time data ingestion/ analysis using Kafka / Spark-streaming.
  • Configured Hive and written Hive UDF's and UDAF's Also, created Static and Dynamic with bucketing as required.
  • Managing and scheduling Jobs on a Hadoop cluster using Oozie.
  • Created Hive External tables and loaded the data in to tables and query data using HQL.
  • Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
  • Developed Oozie workflow for scheduling and orchestrating the ETL process and worked on Oozie workflow engine for job scheduling.
  • Managed and reviewed the Hadoop log files using Shell scripts.
  • Migrated ETL jobs to Pig scripts to do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
  • Supported Senior Engineer in installing and configuring the Hadoop Clusters using Cloudera version CDH4.
  • Using Hive join queries to join multiple tables of a source system and load them to Elastic search tables.
  • Experience in managing and reviewing huge Hadoop log files.
  • Collected the logs data from web servers and integrated in to HDFS using Flume.
  • Expertise in designing and creating various analytical reports and Automated Dashboards to help users to identify critical KPIs and facilitate strategic planning in the organization.
  • Involved in Cluster maintenance, Cluster Monitoring and Troubleshooting
  • Loaded all data-sets into Hive from Source CSV files using spark and Cassandra from Source CSV files using Spark/PySpark.
  • Migrated the computational code in hql to PySpark.
  • Completed data extraction, aggregation and analysis in HDFS by using PySpark and store the data needed to Hive.
  • Developed Python code to gather the data from HBase (Cornerstone) and designs the solution to implement using PySpark.
  • Maintaining technical documentation for each and every step of development environment and launching Hadoop clusters.
  • Good Knowledge on Hadoop Cluster architecture and working with Hadoop clusters using Cloudera (CDH5) and Hortonworks Distributions.
  • Worked on different file formats like Parquette, Orc, Avro, Sequence files using MapReduce/Hive/Impala.
  • Worked with Avro Data Serialization system to work with JSON data formats.
  • Used Amazon Web Services (AWS) S3 to store large amount of data in identical/similar repository.
  • Worked with the Data Science team to gather requirements for various data mining projects.
  • Wrote shell scripts for rolling day-to-day processes and it is automated.
  • Involved in build applications using Maven and integrated with Continuous Integration servers like Jenkins to build jobs.
  • Worked on BI tools as Tableau to create dashboards like weekly, monthly, daily reports using tableau desktop and publish them to HDFS cluster.

Environment: Spark, Spark SQL, Spark Streaming, Scala, Kafka, Hadoop, HDFS, Hive, Oozie, Pig, Nifi, Sqoop, AWS (EC2, S3, EMR), Shell Scripting, HBase, Jenkins, Tableau, Oracle, MySQL, Teradata and AWS.

Confidential - Bedford, MA

Hadoop Developer

Responsibilities:

  • Installed & configured multi-node Hadoop Cluster and performed troubleshooting and monitoring of Hadoop Cluster.
  • Installed and configured Flume, Oozie on the Hadoop cluster.
  • Managing, defining and scheduling Jobs on a Hadoop cluster.
  • Worked on installing cluster, commissioning & decommissioning of datanode, namenode recovery, capacity planning, and slots configuration.
  • Resource management of Hadoop Cluster including adding/removing cluster nodes for maintenance and capacity needs.
  • Wrote Hive queries for ad-hoc reporting and ETL.
  • Worked with different file formats such as Text, Sequence files, Avro, ORC and Parquette.
  • Pushing logs onto Hadoop cluster via SOAP UI calls.
  • Involved in upgrading all the Hadoop components to the latest versions.
  • Worked on NOSQL database Mongo DB.
  • Experience in managing and reviewing Hadoop log files.
  • Optimizing Hadoop MapReduce code, Hive and Pig scripts for better scalability, reliability and performance.
  • Developed the OOZIE workflows for the Application execution.
  • Performing data migration from Legacy Databases RDBMS to HDFS using Sqoop.
  • Worked on Creation of Optimized Environment for testing and demonstrating Spark, Cassandra & Kafka Connectivity with objective of minimizing Spark SQL query times, Over JDBC/ODBC for databases
  • Writing Pig scripts for data processing.
  • Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
  • Implemented Hive tables and HQL Queries for the reports.
  • Imported data from Cassandra into HDFS using Mongo export utility.
  • Involved in developing shell scripts and automated data management from end to end integration work
  • Experience in performing data validation using HIVE dynamic partitioning and bucketing.
  • Written and used complex data type in storing and retrieved data using HQL in Hive.
  • Developed Hive queries to analyze reducer output data.
  • Implemented ETL code to load data from multiple sources into HDFS using pig scripts.
  • Highly involved in designing the next generation data architecture for the unstructured data.
  • Developed PIG Latin scripts to extract data from source system.
  • Created and maintained technical documentation for executing Hive queries and Pig scripts.
  • Involved in Extracting, loading Data from Hive to Load an RDBMS using Sqoop

Environment: Spark, Scala, HDFS, Map Reduce, MySQL, Cassandra, Hive, HBase, Oozie, PIG, ETL, Hortonworks (HDP 2.0), Shell Scripting, Linux, Sqoop, Flume and Oracle 11g.

Confidential - Durham, NC

Sr. Big Data Developer

Responsibilities:

  • As a Sr. Big Data Developer, I worked on Hadoop eco-systems including Hive, HBase, Oozie, Pig, Zookeeper, Spark Streaming MCS (MapR Control System) and so on with MapR distribution.
  • Installed and configured Hadoop MapReduce, HDFS, Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
  • Built code for real time data ingestion using Java, MapR-Streams (Kafka) and STORM.
  • Involved in various phases of development analyzed and developed the system going through Agile Scrum methodology.
  • Worked on Apache Solr which is used as indexing and search engine.
  • Involved in development of Hadoop System and improving multi-node Hadoop Cluster performance.
  • Worked on analyzing Hadoop stack and different Big data tools including Pig and Hive, Hbase database and Sqoop.
  • Developed data pipeline using flume, Sqoop and pig to extract the data from weblogs and store in HDFS
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
  • Worked with different data sources like Avro data files, XML files, JSON files, SQL server and Oracle to load data into Hive tables.
  • Used J2EE design patterns like Factory pattern & Singleton Pattern.
  • Used Spark to create the structured data from large amount of unstructured data from various sources.
  • Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3)
  • Performed transformations, cleaning and filtering on imported data using Hive, MapReduce, Impala and loaded final data into HDFS.
  • Developed Python scripts to find vulnerabilities with SQL Queries by doing SQL injection.
  • Experienced in designing and developing POC’s in Spark using Scala to compare the performance of Spark with Hive and SQL/Oracle.
  • Responsible for coding MapReduce program, Hive queries, testing and debugging the MapReduce programs.
  • Extracted Real time feed using Spark streaming and convert it to RDD and process data into Data Frame and load the data into Cassandra.
  • Involved in the process of data acquisition, data pre-processing and data exploration of telecommunication project in Scala.
  • Implemented a distributed messaging queue to integrate with Cassandra using Apache Kafka and Zookeeper.
  • Specified the cluster size, allocating Resource pool, Distribution of Hadoop by writing the specification texts in JSON File format.
  • Imported weblogs & unstructured data using the Apache Flume and stores the data in Flume channel.
  • Exported event weblogs to HDFS by creating a HDFS sink which directly deposits the weblogs in HDFS.
  • Used RESTful web services with MVC for parsing and processing XML data.
  • Utilized XML and XSL Transformation for dynamic web-content and database connectivity.
  • Involved in loading data from Unix file system to HDFS. Involved in designing schema, writing CQL's and loading data using Cassandra.
  • Built the automated build and deployment framework using Jenkins, Maven etc.

Environment: Hadoop, Hive, HBase, Oozie, Pig, Zookeeper, MapR, HDFS, MapReduce, Java, MS, Jenkins, Agile, Apache Solr, Apache Flume, Amazon EMR, Spark, Scala, Cassandra, Apache Kafka, MVC

Confidential - Hillsboro, OR

Sr. Java/Hadoop Developer

Responsibilities:

  • Massively involved in software development life cycle starting from gathering requirements and performing Object Oriented Analysis.
  • Designed and developed various modules of the application with J2EE design architecture, Spring MVC architecture and Spring Bean Factory using IOC, AOP concepts.
  • Used MAVEN for developing build scripts and deploying the application onto WebLogic.
  • Implemented Spark RDD transformations to map business analysis and apply actions on top of transformations.
  • Developed Soap and REST service using Spring MVC framework for all the modules.
  • Participated in the dynamic form generation, automatic completion of forms and functions of user validation using Ajax.
  • Exported data from DB2 to HDFS using Sqoop and Developed MapReduce jobs using Java API.
  • Designed and Developed application modules using spring and Hibernate frameworks.
  • Utilized Core J2EE designs patterns such as Singleton in the implementation of the services.
  • Implemented MVC architecture by separating the business logic from the presentation layer using spring.
  • Worked on JDBC framework encapsulated using DAO pattern to connect to the database.
  • Worked with Maven for build scripts and Setup the Log4J Logging framework.
  • Implemented MapReduce programs to handle semi/unstructured data like XML, JSON, and sequence files for log files.
  • Worked on analyzing, writing Hadoop MapReduce jobs using Java API, Pig and Hive.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Used Maven for developing build scripts and deploying the application onto WebLogic.
  • Involved in configuring builds using Jenkins with Git and used Jenkins to deploy the applications onto Dev, QA environments.
  • Worked with JSON objects and JavaScript and JQuery intensively to create interactive web pages.
  • Involved in unit testing, system integration testing and enterprise user testing using JUnit.
  • Involved in Setup and benchmark of Hadoop /HBase clusters for internal use.
  • Developed JSP and Java classes for various transactional/ non-transactional reports of the system using extensive SQL queries.
  • Developed the UI Screens using JSP and HTML and did the client side validation with the JavaScript.

Environment: J2EE, Spring, MVC, Ajax, DB2, HDFS, Sqoop, MapReduce, Java, Hibernate, Maven, XML, JSON, Hadoop, Hive, Pig, Jenkins, JavaScript, JQuery, HBase, JUnit

Confidential

Java Developer

Responsibilities:

  • Developed software test plans, test design specifications, and test script for various test scenarios.
  • Responsible for the verification of the SOLR search and indexes working and the quality before it is published.
  • SOLR tuning with various search strategies in the customer in-house developed data sets.
  • Responsible for developing DAO layer using Spring MVC and configuration XML for Hibernate.
  • Wrote Hibernate classes, DAO's to retrieve & store data, configured Hibernate files.
  • Used spring framework for dependency injection, transaction management.
  • Used Spring MVC framework controllers for Controllers part of the MVC.
  • Involved in developing the UI pages using HTML, DHTML, JavaScript, Ajax, JQuery, JSP and tag libraries.
  • Extensively used Java Collection framework and Exception handling.
  • Styling in CSS and JSPs is done as per the Style guide provided by UI team.
  • Used JavaScript for client side validations. Used JUnit for unit testing of the system and Log4J for logging.
  • Extensively used Eclipse IDE for developing, debugging, integrating and deploying the application.
  • Developed the presentation layer using JSP, HTML and client side validations using JavaScript.
  • Wrote and debugged the Maven Scripts for building the entire web application.
  • Developed an application using core and advanced java along the PL/SQL Database.
  • Designed Presentation layer using spring framework, JSP and did front-end validations using JavaScript and JQuery.
  • Involved in design and development of UI component, using frameworks Angular JS, Ember JS, JavaScript, HTML, CSS and Bootstrap.
  • Designed and developed Ajax calls to populate screens parts on demand.
  • Developed user interface using JSP, JSP Tag libraries and Struts Tag Libraries to simplify the complexities of the application.

Environment: SOLR, Spring, XML, Hibernate, MVC, JavaScript, Ajax, JQuery, Java, CSS, JUnit, HTML, Bootstrap, Angular JS, Eclipse, JSP, Maven, Angular JS, CSS

Confidential

Java Developer

Responsibilities:

  • Developed the use cases and class diagrams using Rational Rose/UML.
  • Performed end-to-end design and development of all layers of the application.
  • Implemented Spring MVC for designing and implementing the UI Layer for the application.
  • Wrote Spring Validator, Spring AOP for validating the input data.
  • Used Hibernate ORM in the persistence layer and implemented DAO’s to access data from with Oracle and MYSQL databases.
  • Used XML, WSDL, UDDI and SOAP Web Services (JAX-WS) using Apache Axis2 framework for communicating data between different applications.
  • Storing the SOAP messages received in the JMS Queue of WebSphere MQ (MQ Series).
  • Develop JAX-WS services and JSR-286 compliant Portlet using Java Server Faces (JSF).
  • Developed Data access bean and developed EJB s that are used to access data from the database.
  • Used EJB to inject the services and their dependencies.
  • Involved in Coding HTML, CSS, JavaScript for UI validation for dynamic manipulation of the elements on the screen and to validate the input.
  • Wrote PL/SQL and SQL blocks for the application.
  • Used Core java Multi-Threading concepts for avoiding concurrent processes.
  • Tested all the components in application using Junit framework.
  • Responsible for deploying application file on IBM WebSphere Application server.
  • Used Log4j package for logging, ANT for automated deployment and Junit for Testing.

Environment: J2EE, JDK, Spring MVC 3.x, Spring AOP 3.x, EJB 1.x, Java Beans, SOAP Web Services, Apache-Axis1, JSR -286 Portlet, JSF 2.x, JMS, Hibernate, JSP, XML, JNDI, Design Patterns, TOAD, IBM WebSphere, Junit, ANT, PL/SQL, Oracle 9i, MYSQL, Rational Rose, Unix.

We'd love your feedback!