Hadoop Developer Resume
Atlanta, Ga
SUMMARY:
- Around 8+ years of programming experience involved in all phases of Software Development Life Cycle (SDLC).
- Around 4 Years of Big data related architecture experience developing Hadoop applications.
- Experience with configuration of Hadoop Ecosystem components: MapReduce, Pig, Hive, Sqoop, HBase, Flume, Oozie.
- Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, HBase, Hive, MapReduce and Sqoop.
- Implemented POC to migrate map reduce jobs into Spark RDD transformations using Scala.
- Developed Apache Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying.
- Experience with performing real time analytics on NoSQL data bases like HBase and Cassandra .
- Good knowledge in working with Impala, Storm and Kafka.
- Experience with Dimensional modeling, Data migration, Data cleansing, Data profiling, and ETL Processes features for data warehouses.
- Worked with Oozie work flow engine to schedule time based jobs to perform multiple actions.
- Experience in importing and exporting data from RDBMS into HDFS using Sqoop.
- Hands on experience in working with database like Oracle, MySQL and PL/SQL.
- Experience in high level ETL architecture for overall data transfer from the OLTP to OLAP
- Good working knowledge on processing Batch applications.
- Experience in writing MapReduce programs and UDFs for both Hive and Pig in Java.
- Used Flume to channel data from different sources to HDFS.
- Experience with Maven, Jenkins and GIT.
- Good experience in Hive partitioning, bucketing and perform different types of joins on Hive tables and implementing Hive SerDe like JSON and Avro.
- Experience in using different file formats - Avro, Sequence Files, ORC, JSON and Parquet.
- Experience in Performance Tuning, Optimization and Customization.
- Experience with Unit Testing Map Reduce programs using MRUnit, JUnit.
- Experience in Active Development as well as onsite coordination activities in web-based, client/server and distributed architecture using Java, J2EE which includes Web services, Spring, Struts, Hibernate and JSP/Servlets along with incorporating MVC architecture.
- Good working knowledge on servers like Tomcat, Web Logic 8.0.
- Ability to work in teams as well as an individual, quick learner and able to meet deadlines.
TECHNICAL SKILLS:
Hadoop 2.0: Hadoop, CDH5.3.2, MapReduce, YARN, Spark 1.4, Sqoop, Hive, Oozie, PIG, HDFS, Flume, Impala, Storm 0.9.5, Apace Kafka 0.9.0.0
Programming Languages: C, Java 6.0, Scala 2.11 MySQL, PL/SQL, PIG Latin, HiveQL and Python
Web Technologies: HTML, CSS, JavaScript, Ajax, XML
Java &J2EE Technologies: Core Java, Servlets, JSP, JDBC, EJB’s
Databases: Hive QL, HBase, MongoDB, Cassandra 2.2.0,
IDE: Eclipse
Application Server: Apache Tomcat 7.0, Apache Tomcat 6.0, WebSphere
Process Automation Tools: SVN, JUnit
Design Patterns: MVC, Singleton, Factory
Frameworks: Spring 3.0.5, Hibernate 3.5.1, Struts 1.3.10, EJB, JUnit, MRUnit
PROFESSIONAL EXPERIENCE:
Confidential, Atlanta, GA
Hadoop Developer
Responsibilities:
- Assisted in the development of high quality analytical reports of weather and GIS data.
- Used Talend to generate optimized code to load, transform, enrich, and cleanse data inside Hadoop.
- Imported unstructured data like logs from different web servers to HDFS using Flume and developed MapReduce jobs for log analysis, recommendations and analytics.
- Involved in real-time data processing using Storm.
- Expertise in real-time analytics, machine learning and continuous monitoring of operations using Storm.
- Converted and loaded local data files into HDFS through the Unix shell.
- Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
- Created HBase tables to store variable data formats coming from different portfolios.
- Implemented HBase Co-processors to notify Support team when inserting data into HBase Tables.
- Used Kafka along with HBase to render data streaming.
- Implemented Spark for fast interactive data analysis of datasets loaded in RDD.
- Analyzed the data by performing Hive queries (HiveQL), ran Pig scripts, Spark SQL and Spark streaming.
- Used Hive for data warehousing and summarization.
- Used Pig for aggregation, cleaning and incremental ETL functions and developing UDFs for filtering.
- Developed Pig UDFs to specifically preprocess and filter data sets for analysis.
- Continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
- Wrote custom MapReduce jobs for cleaning and reorganizing data, accounting for and removing statistical anomalies.
- Used Teradata database management system to manage the warehousing operations and parallel processing.
- Validated data sets graphically with Excel and did touch ups in Photoshop.
- Designed & scheduled workflows for updating system reports using Oozie.
Environment: Cloudera, Avro, HBase, HDFS, Hive, Pig, Java, SQL, Sqoop, Flume, Oozie, Java (jdk 1.7), Eclipse, YARN, SQL Server, Spark, Zookeeper, SVN, Talend
Confidential, CA
Hadoop Developer
Responsibilities:
- Worked on importing data from various sources and performed transformations using MapReduce, Hive to load data into HDFS.
- Configured Sqoop jobs to import data from RDBMS into HDFS using Oozie workflows.
- Worked on setting up Pig, Hive and HBase on multiple nodes and developed using Pig, Hive, HBase and MapReduce.
- Solved small file problem using Sequence files processing in Map Reduce.
- Written various Hive and Pig scripts.
- Created HBase tables to store variable data formats coming from different portfolios.
- Performed real time analytics on HBase using Java API and Rest API.
- Implemented HBase Co-processors to notify Support team when inserting data into HBase Tables.
- Worked on compression mechanisms to optimize MapReduce Jobs.
- Analyzed the customer behavior by performing click stream analysis and to ingest the data used flume.
- Experienced with working on Avro Data files using Avro Serialization system.
- Implemented business logic by writing UDF's in Java and used various UDF's from Piggybanks and other sources.
- Continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
- Unit tested and tuned SQLs and ETL Code for better performance.
- Monitored the performance and identified performance bottlenecks in ETL code.
- Used TABLEAU which grabs data to generate reports, graphs and charts summarizing the given set of data.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
Environment: Horton works, Map Reduce, HBase, HDFS, Hive, Pig, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Java, Eclipse
Confidential,San Francisco, CA
Big Data/Hadoop Developer
Responsibilities:
- Created Hive tables and working on them using Hive QL.
- Involved in installing Hadoop Ecosystem components.
- Validated Name node, Data node status in a HDFS cluster.
- Importing and exporting data from HDFS to RDBMS and vice-versa using SQOOP.
- Experienced in developing HIVE Queries on different data formats like Text file, CSV file.
- Developed multiple MapReduce jobs in java for data cleaning and preprocessing.
- Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis
- Installed and configured Hadoop cluster in Test and Production environments
- Performed both major and minor upgrades to the existing CDH cluster
- Code review as per the customer coding standards.
- Testing and providing the valid test data to users as per requirement.
- Weekly meetings with technical collaborators and active participation in code review sessions with senior and junior developers.
- Responsible to manage data coming from different sources.
- Supporting Hbase Architecture Design with the Hadoop Architect team to develop a Database Design in HDFS.
- Involved in HDFS maintenance and loading of structured and unstructured data.
- Wrote Hive queries for data analysis to meet the business requirements.
- Installed and configured Pig and also written Pig Latin scripts.
- Developed UDFs for Pig Data Analysis.
- Involved in managing and reviewing Hadoop log files.
- Developed Scripts and Batch Job to schedule various Hadoop Program.
- Utilized Agile Scrum Methodology to help manage and organize a team of 4 developers with regular code review sessions.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce.
- Analyzed the data by performing Hive queries and running Pig scripts to know user behavior.
- Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
- Installed Oozie workflow engine to run multiple Hive and Pig jobs.
- Developed Hive queries to process the data and generate the data cubes for visualizing.
Environment: Java, Hadoop, MapReduce, HDFS, Hive, Pig, Sqoop, Linux, XML, Eclipse, Cloudera
Confidential, Warren, NJ
Java Developer
Project Abstract:
T he AgentPortal project involves the design and development of a Web based application to provide online insurance services. The web application provides functionalities such as policy management, profile management, and claim reporting. As a Java developer, my responsibility was to develop enterprise integration solutions utilizing WebSphere Application Server/Java. In addition, designing Web-based applications, Web services and MQ/JMS services solutions based on business need and maintaining appropriate systems documentation was also my liability.
Responsibilities:
- Involved in study of User Requirement Specification, Communicated with Business Analysts to resolve ambiguity in Requirements document.
- Worked in Agile Scrum Methodology.
- Deployed and Developed Java/J2EE based application
- Developed application using Spring MVC and Hibernate as the ORM tool.
- Developed user interface with HTML, CSS, JavaScript, JQuery, AJAX and JSP
- Connected with database using MYSQL and PL/SQL
- Created SOAP based webservice to provide the service of Location Look up for the agents.
- Created a Subscriber-Publisher Messaging Topics using JMS to send and receive messages.
- Created Daemon threads for certain long running reporting process.
- Scheduled a certain tasks using Java Timer Class.
- Interfaced with Oracle back-end using Hibernate Framework.
- Implemented Test Driven Development using JUnit.
- Deployed applications on WebSphere during development
- Wrote shell scripts to schedule standalone Java programs.
- The message transport mechanism supported various data formats like XML and JSON.
Environment- Java, Spring, Hibernate, JSP, HTML, CSS, XML, JavaScript, JQuery, JUnit, AJAX, Multi-Threading, Oracle, Web Service - SOAP, WebSphere, MYSQL.
Confidential, Philadelphia, PA
Intern Java Developer
Responsibilities:
- Web application development using Java
- Developed a high-end presentation layer with CSS, JavaScript, Ajax and JSP
- Expertise as JavaScript developer in OOP patterns
- Worked to build web application using DB2
- Wrote the middleware tool using Spring framework
- Used IBM tool Integration Developer 8.5 to develop the application.
- Performed Unit testing of the application.
- Used message broker tool for testing.
- Developed components using session and deployed them on Websphere 7.0 environment
- Used Hibernate as ORM tool for persisting Java project to relation in database
- Hands on experience using Agile software development methodologies
- Developed and performed effective unit testing and continuous integration
- Involved in technical documentations
Environment: JSP, CSS, JavaScript, AJAX, XML, DB2, Websphere 7.0, Spring, Eclipse, Hibernate, Agile, Integration Developer, Oracle, SVN
