Hadoop Developer Resume
Woonsocket, RI
PROFESSIONAL SUMMARY:
- Over 9+ years of experience in IT industry with 4 years of experience in Big Data implementing complete Hadoop solutions along with Scala and Java/J2EE.
- Good working experience in using Apache Hadoop eco system components like MapReduce, HDFS, Hive, Sqoop, Pig, Oozie, Flume, HBase and Zoo Keeper.
- Extensive experience working on spark in performing ETL using Spark Core, Spark - SQL and Real-time data processing using Spark Streaming.
- Expertise in working with RDD’s, DataFrames and DataSet API’s.
- Extensively worked with Kafka as middleware for real-time data pipelines.
- Writing UDFs and integrating with Hive and Pig using Java.
- Experience with Sequence files, AVRO and ORC file formats and compression.
- Experience in Hadoop Distributions: Cloudera and Hortonworks,
- Performed importing and exporting data into HDFS and Hive using Sqoop.
- Experience in job workflow scheduling tools like Oozie.
- Extensive knowledge in using SQL Queries for backend database analysis.
- Strong knowledge in NOSQL column oriented databases like HBase, Cassandra and its integration with Hadoop cluster.
- Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems (RDBMS) and vice-versa.
- Led many Data Analysis & Integration efforts involving HADOOP along with ETL.
- Hands on experience on Enterprise Data Lake to provide support for various uses cases including Analytics, processing, storing and Reporting of voluminous, rapidly changing, structured and unstructured data.
- Extensive experience with SQL, PL/SQL and database concepts.
- Transferred bulk data from RDBMS systems like Teradata into HDFS using Sqoop.
- Experience in installation, configuration, supporting and managing Hadoop Clusters using Horton works, and Cloudera (CDH3, CDH4) distributions on Amazon web services (AWS)
- Experience in analyzing data using Hive QL, Pig Latin, and custom MapReduce programs in Java.
- Strong experience in analyzing large amounts of data sets writing PySpark scripts and Hive queries.
- Worked in developing a Nifi flow prototype for data ingestion in HDFS.
- Well-versed in Agile, other SDLC methodologies and can coordinate with owners and SMEs.
- Worked on different operating systems like UNIX, Linux, and Windows
- Diverse experience utilizing Java tools in business, Web, and client-server environments including Java Platform, Enterprise Edition (Java EE), Enterprise Java Bean (EJB), JavaServer Pages (JSP), Java Servlets (including JNDI), Struts, and Java database Connectivity (JDBC) technologies.
- Fluid understanding of multiple programming languages, including C#, C, C++, JavaScript, HTML, and XML.
- Experience in web application design using open source MVC , Spring and Struts Frameworks.
TECHNICAL SKILLS:
Hadoop Core Services: HDFS, Map Reduce, Spark, YARN
Hadoop Distribution: Horton works, Cloudera, Apache
NO SQL Databases: HBase, Cassandra
Hadoop Data Services: Hive, Pig, Sqoop, Flume
Hadoop Operational Services: Zookeeper, Oozie
Monitoring Tools: Cloudera Manager
Cloud Computing Tools: Amazon AWS
Languages: C, Java/J2EE, Python, SQL, PL/SQL, Pig Latin, HiveQL, Unix Shell Scripting
Java & J2EE Technologies: Core Java, Servlets, Hibernate, Spring, Struts
Application Servers: Web Logic, Web Sphere, JBoss, Tomcat
Databases: Oracle, MySQL, Postgress, Teradata
Operating Systems: UNIX, Windows, LINUX
Build Tools: Jenkins, Maven, ANT
Development Tools: Microsoft SQL Studio, Toad, Eclipse, NetBeans
Development methodologies: Agile/Scrum
Visualization and analytics too: Tableau Software, Qlik View
PROFESSIONAL EXPERIENCE:
Confidential - Woonsocket, RI
Hadoop Developer
Responsibilities:
- Imported the data from various formats like JSON, Sequential, Text, CSV, AVRO and Parquet to HDFS cluster with compressed for optimization.
- Worked on ingesting data from RDBMS sources like - Oracle, SQL Server and Teradata into HDFS using Sqoop.
- Loaded all data-sets into Hive from Source CSV files using spark and Cassandra from Source CSV files using Spark
- Built a Ingestion Framework that would ingest the files from SFTP to HDFS using Apache NIFI and ingest Financial data into HDFS
- Created environment to access Loaded Data via spark SQL, through JDBC&ODBC (via Spark Thrift Server). Developed real time data ingestion/ analysis using Kafka / Spark-streaming.
- Configured Hive and written Hive UDF's and UDAF's Also, created Static and Dynamic with bucketing as required.
- Managing and scheduling Jobs on a Hadoop cluster using Oozie.
- Created Hive External tables and loaded the data in to tables and query data using HQL.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Developed Oozie workflow for scheduling and orchestrating the ETL process and worked on Oozie workflow engine for job scheduling.
- Managed and reviewed the Hadoop log files using Shell scripts.
- Migrated ETL jobs to Pig scripts to do Transformations, even joins and some pre-aggregations before storing the data onto HDFS.
- Supported Senior Engineer in installing and configuring the Hadoop Clusters using Cloudera version CDH4.
- Using Hive join queries to join multiple tables of a source system and load them to Elastic search tables.
- Experience in managing and reviewing huge Hadoop log files.
- Collected the logs data from web servers and integrated in to HDFS using Flume.
- Expertise in designing and creating various analytical reports and Automated Dashboards to help users to identify critical KPIs and facilitate strategic planning in the organization.
- Involved in Cluster maintenance, Cluster Monitoring and Troubleshooting
- Loaded all data-sets into Hive from Source CSV files using spark and Cassandra from Source CSV files using Spark/PySpark.
- Migrated the computational code in hql to PySpark.
- Completed data extraction, aggregation and analysis in HDFS by using PySpark and store the data needed to Hive.
- Developed Python code to gather the data from HBase (Cornerstone) and designs the solution to implement using PySpark.
- Maintaining technical documentation for each and every step of development environment and launching Hadoop clusters.
- Good Knowledge on Hadoop Cluster architecture and working with Hadoop clusters using Cloudera (CDH5) and Hortonworks Distributions.
- Worked on different file formats like Parquette, Orc, Avro, Sequence files using MapReduce/Hive/Impala.
- Worked with Avro Data Serialization system to work with JSON data formats.
- Used Amazon Web Services (AWS) S3 to store large amount of data in identical/similar repository.
- Worked with the Data Science team to gather requirements for various data mining projects.
- Wrote shell scripts for rolling day-to-day processes and it is automated.
- Involved in build applications using Maven and integrated with Continuous Integration servers like Jenkins to build jobs.
- Worked on BI tools as Tableau to create dashboards like weekly, monthly, daily reports using tableau desktop and publish them to HDFS cluster.
Environment: Spark, Spark SQL, Spark Streaming, Scala, Kafka, Hadoop, HDFS, Hive, Oozie, Pig, Nifi, Sqoop, AWS (EC2, S3, EMR), Shell Scripting, HBase, Jenkins, Tableau, Oracle, MySQL, Teradata and AWS.
Confidential - Bedford, MA
Hadoop Developer
Responsibilities:
- Installed & configured multi-node Hadoop Cluster and performed troubleshooting and monitoring of Hadoop Cluster.
- Installed and configured Flume, Oozie on the Hadoop cluster.
- Managing, defining and scheduling Jobs on a Hadoop cluster.
- Worked on installing cluster, commissioning & decommissioning of datanode, namenode recovery, capacity planning, and slots configuration.
- Resource management of Hadoop Cluster including adding/removing cluster nodes for maintenance and capacity needs.
- Wrote Hive queries for ad-hoc reporting and ETL.
- Worked with different file formats such as Text, Sequence files, Avro, ORC and Parquette.
- Pushing logs onto Hadoop cluster via SOAP UI calls.
- Involved in upgrading all the Hadoop components to the latest versions.
- Worked on NOSQL database Mongo DB.
- Experience in managing and reviewing Hadoop log files.
- Optimizing Hadoop MapReduce code, Hive and Pig scripts for better scalability, reliability and performance.
- Developed the OOZIE workflows for the Application execution.
- Performing data migration from Legacy Databases RDBMS to HDFS using Sqoop.
- Worked on Creation of Optimized Environment for testing and demonstrating Spark, Cassandra & Kafka Connectivity with objective of minimizing Spark SQL query times, Over JDBC/ODBC for databases
- Writing Pig scripts for data processing.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Implemented Hive tables and HQL Queries for the reports.
- Imported data from Cassandra into HDFS using Mongo export utility.
- Involved in developing shell scripts and automated data management from end to end integration work
- Experience in performing data validation using HIVE dynamic partitioning and bucketing.
- Written and used complex data type in storing and retrieved data using HQL in Hive.
- Developed Hive queries to analyze reducer output data.
- Implemented ETL code to load data from multiple sources into HDFS using pig scripts.
- Highly involved in designing the next generation data architecture for the unstructured data.
- Developed PIG Latin scripts to extract data from source system.
- Created and maintained technical documentation for executing Hive queries and Pig scripts.
- Involved in Extracting, loading Data from Hive to Load an RDBMS using Sqoop
Environment: Spark, Scala, HDFS, Map Reduce, MySQL, Cassandra, Hive, HBase, Oozie, PIG, ETL, Hortonworks (HDP 2.0), Shell Scripting, Linux, Sqoop, Flume and Oracle 11g.
Confidential - Durham, NC
Sr. Big Data Developer
Responsibilities:
- As a Sr. Big Data Developer, I worked on Hadoop eco-systems including Hive, HBase, Oozie, Pig, Zookeeper, Spark Streaming MCS (MapR Control System) and so on with MapR distribution.
- Installed and configured Hadoop MapReduce, HDFS, Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
- Built code for real time data ingestion using Java, MapR-Streams (Kafka) and STORM.
- Involved in various phases of development analyzed and developed the system going through Agile Scrum methodology.
- Worked on Apache Solr which is used as indexing and search engine.
- Involved in development of Hadoop System and improving multi-node Hadoop Cluster performance.
- Worked on analyzing Hadoop stack and different Big data tools including Pig and Hive, Hbase database and Sqoop.
- Developed data pipeline using flume, Sqoop and pig to extract the data from weblogs and store in HDFS
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Worked with different data sources like Avro data files, XML files, JSON files, SQL server and Oracle to load data into Hive tables.
- Used J2EE design patterns like Factory pattern & Singleton Pattern.
- Used Spark to create the structured data from large amount of unstructured data from various sources.
- Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3)
- Performed transformations, cleaning and filtering on imported data using Hive, MapReduce, Impala and loaded final data into HDFS.
- Developed Python scripts to find vulnerabilities with SQL Queries by doing SQL injection.
- Experienced in designing and developing POC’s in Spark using Scala to compare the performance of Spark with Hive and SQL/Oracle.
- Responsible for coding MapReduce program, Hive queries, testing and debugging the MapReduce programs.
- Extracted Real time feed using Spark streaming and convert it to RDD and process data into Data Frame and load the data into Cassandra.
- Involved in the process of data acquisition, data pre-processing and data exploration of telecommunication project in Scala.
- Implemented a distributed messaging queue to integrate with Cassandra using Apache Kafka and Zookeeper.
- Specified the cluster size, allocating Resource pool, Distribution of Hadoop by writing the specification texts in JSON File format.
- Imported weblogs & unstructured data using the Apache Flume and stores the data in Flume channel.
- Exported event weblogs to HDFS by creating a HDFS sink which directly deposits the weblogs in HDFS.
- Used RESTful web services with MVC for parsing and processing XML data.
- Utilized XML and XSL Transformation for dynamic web-content and database connectivity.
- Involved in loading data from Unix file system to HDFS. Involved in designing schema, writing CQL's and loading data using Cassandra.
- Built the automated build and deployment framework using Jenkins, Maven etc.
Environment: Hadoop, Hive, HBase, Oozie, Pig, Zookeeper, MapR, HDFS, MapReduce, Java, MS, Jenkins, Agile, Apache Solr, Apache Flume, Amazon EMR, Spark, Scala, Cassandra, Apache Kafka, MVC
Confidential - Hillsboro, OR
Sr. Java/Hadoop Developer
Responsibilities:
- Massively involved in software development life cycle starting from gathering requirements and performing Object Oriented Analysis.
- Designed and developed various modules of the application with J2EE design architecture, Spring MVC architecture and Spring Bean Factory using IOC, AOP concepts.
- Used MAVEN for developing build scripts and deploying the application onto WebLogic.
- Implemented Spark RDD transformations to map business analysis and apply actions on top of transformations.
- Developed Soap and REST service using Spring MVC framework for all the modules.
- Participated in the dynamic form generation, automatic completion of forms and functions of user validation using Ajax.
- Exported data from DB2 to HDFS using Sqoop and Developed MapReduce jobs using Java API.
- Designed and Developed application modules using spring and Hibernate frameworks.
- Utilized Core J2EE designs patterns such as Singleton in the implementation of the services.
- Implemented MVC architecture by separating the business logic from the presentation layer using spring.
- Worked on JDBC framework encapsulated using DAO pattern to connect to the database.
- Worked with Maven for build scripts and Setup the Log4J Logging framework.
- Implemented MapReduce programs to handle semi/unstructured data like XML, JSON, and sequence files for log files.
- Worked on analyzing, writing Hadoop MapReduce jobs using Java API, Pig and Hive.
- Responsible for building scalable distributed data solutions using Hadoop.
- Used Maven for developing build scripts and deploying the application onto WebLogic.
- Involved in configuring builds using Jenkins with Git and used Jenkins to deploy the applications onto Dev, QA environments.
- Worked with JSON objects and JavaScript and JQuery intensively to create interactive web pages.
- Involved in unit testing, system integration testing and enterprise user testing using JUnit.
- Involved in Setup and benchmark of Hadoop /HBase clusters for internal use.
- Developed JSP and Java classes for various transactional/ non-transactional reports of the system using extensive SQL queries.
- Developed the UI Screens using JSP and HTML and did the client side validation with the JavaScript.
Environment: J2EE, Spring, MVC, Ajax, DB2, HDFS, Sqoop, MapReduce, Java, Hibernate, Maven, XML, JSON, Hadoop, Hive, Pig, Jenkins, JavaScript, JQuery, HBase, JUnit
Confidential
Java Developer
Responsibilities:
- Developed software test plans, test design specifications, and test script for various test scenarios.
- Responsible for the verification of the SOLR search and indexes working and the quality before it is published.
- SOLR tuning with various search strategies in the customer in-house developed data sets.
- Responsible for developing DAO layer using Spring MVC and configuration XML for Hibernate.
- Wrote Hibernate classes, DAO's to retrieve & store data, configured Hibernate files.
- Used spring framework for dependency injection, transaction management.
- Used Spring MVC framework controllers for Controllers part of the MVC.
- Involved in developing the UI pages using HTML, DHTML, JavaScript, Ajax, JQuery, JSP and tag libraries.
- Extensively used Java Collection framework and Exception handling.
- Styling in CSS and JSPs is done as per the Style guide provided by UI team.
- Used JavaScript for client side validations. Used JUnit for unit testing of the system and Log4J for logging.
- Extensively used Eclipse IDE for developing, debugging, integrating and deploying the application.
- Developed the presentation layer using JSP, HTML and client side validations using JavaScript.
- Wrote and debugged the Maven Scripts for building the entire web application.
- Developed an application using core and advanced java along the PL/SQL Database.
- Designed Presentation layer using spring framework, JSP and did front-end validations using JavaScript and JQuery.
- Involved in design and development of UI component, using frameworks Angular JS, Ember JS, JavaScript, HTML, CSS and Bootstrap.
- Designed and developed Ajax calls to populate screens parts on demand.
- Developed user interface using JSP, JSP Tag libraries and Struts Tag Libraries to simplify the complexities of the application.
Environment: SOLR, Spring, XML, Hibernate, MVC, JavaScript, Ajax, JQuery, Java, CSS, JUnit, HTML, Bootstrap, Angular JS, Eclipse, JSP, Maven, Angular JS, CSS
Confidential
Java Developer
Responsibilities:
- Developed the use cases and class diagrams using Rational Rose/UML.
- Performed end-to-end design and development of all layers of the application.
- Implemented Spring MVC for designing and implementing the UI Layer for the application.
- Wrote Spring Validator, Spring AOP for validating the input data.
- Used Hibernate ORM in the persistence layer and implemented DAO’s to access data from with Oracle and MYSQL databases.
- Used XML, WSDL, UDDI and SOAP Web Services (JAX-WS) using Apache Axis2 framework for communicating data between different applications.
- Storing the SOAP messages received in the JMS Queue of WebSphere MQ (MQ Series).
- Develop JAX-WS services and JSR-286 compliant Portlet using Java Server Faces (JSF).
- Developed Data access bean and developed EJB s that are used to access data from the database.
- Used EJB to inject the services and their dependencies.
- Involved in Coding HTML, CSS, JavaScript for UI validation for dynamic manipulation of the elements on the screen and to validate the input.
- Wrote PL/SQL and SQL blocks for the application.
- Used Core java Multi-Threading concepts for avoiding concurrent processes.
- Tested all the components in application using Junit framework.
- Responsible for deploying application file on IBM WebSphere Application server.
- Used Log4j package for logging, ANT for automated deployment and Junit for Testing.
Environment: J2EE, JDK, Spring MVC 3.x, Spring AOP 3.x, EJB 1.x, Java Beans, SOAP Web Services, Apache-Axis1, JSR -286 Portlet, JSF 2.x, JMS, Hibernate, JSP, XML, JNDI, Design Patterns, TOAD, IBM WebSphere, Junit, ANT, PL/SQL, Oracle 9i, MYSQL, Rational Rose, Unix.
