We provide IT Staff Augmentation Services!

Sr. Hadoop Developer Resume

0/5 (Submit Your Rating)

Sunnyvale, CA

SUMMARY

  • Around 9 years of experience in Application analysis, Design, Development, Maintenance and Supporting web, Client - server based applications in Java/J2EE technologies which includes 3 years of experience with Big Data and Hadoop related components like HDFS, Map Reduce, Pig, Hive, YARN, Sqoop, Flume, Crunch, Spark, Strom, Scala, Kafka.
  • Experience in multiple Hadoop distributions like Cloudera, MapR, and Horton works.
  • Excellent understanding of NoSQL databases like HBase, Cassandra and MongoDB.
  • Experience on working structured, unstructured data with various file formats such as Avro data files, xml files, JSON files, sequence files using Map Reduce programs.
  • Work experience with cloud configurations like Amazon web services (AWS).
  • Implemented custom business logic and performed join optimization, secondary sorting, custom sorting using Map Reduce programs.
  • Experienced testing and running of Map Reduce pipelines on Apache Crunch.
  • Expertise in Data ingestion using SQOOP, Apache Kafka, Spark Streaming and FLUME
  • Implemented business logic using Pig scripts. Wrote custom Pig UDF’s to analyze data
  • Performed different ETL operations using Pig for joining operations and transformations on data to join, clean, aggregate and analyze data.
  • Experience with Oozie Workflow Engine to automate and parallelize Hadoop, Map Reduce and Pig jobs
  • Extensive experience with wiring SQL queries using HiveQL to perform analytics on structured data.
  • Experience in performing data validation using HIVE dynamic partitioning and bucketing.
  • Experienced in importing and exporting data between RDBMS and Teradata into HDFS using Sqoop.
  • Experienced in handling streaming data like web server log data using flume.
  • Worked on Cassandra database and related web services for storing unstructured data.
  • Good noledge analyzing data using Python development and scripting for Hadoop Streaming.
  • Experience in implementing algorithms for analyzing using spark.
  • Experience in implementing Spark using Scala and SparkSQL for faster processing of data.
  • Experience in getting data from various sources into HDFS and building reports using Tableau.
  • Experience in creating tables on top of data on AWS S3 obtained from different data sources and providing them to analytics team building reports using Tableau.
  • Extensive Hands on experience with Accessing and perform CURD operations against HBase data using Java API and implementing time series data management.
  • Involved in various data mining tasks such as pattern mining, classification and clustering techniques.
  • Expert noledge over J2EE Design Patterns like MVC Architecture, Singleton, Factory Pattern, Front Controller, Session Facade, Business Delegate and Data Access Object for building J2EE Applications.
  • Experienced in J2EE, Spring, Hibernate, SOAP/Rest web services,JMS, JNDI, EJB, JAX-WS.
  • Expertise with Application servers and web servers like Oracle WebLogic, IBM WebSphere, Apache Tomcat, JBOSS and VMware.
  • Proven expertise in implementing IOC/Dependency Injection features in various aspects of Spring Framework.
  • Experienced in developing the unit test cases using MRUnit, JUnit and Easy Mock.
  • Knowledge on Splunk for logging mechanism.
  • Knowledge on Build tool Jenkins.
  • Experience in using Maven and ANT for build automation.
  • Experience in using version control and configuration management tools like SVN, CVS.
  • Experience working in environments using Agile (SCRUM) and Waterfall methodologies.
  • Experience in designing applications using UML Diagrams like Class Diagram, Component Diagram, Sequence Diagrams, and Deployment Diagram using MS Visio, Rational Rose.
  • Expertise in database modeling, administration and development usingSQL and PL/SQL in Oracle (8i, 9i and 10g), MySQL, Teradata, DB2and SQL Server environments.

TECHNICAL SKILLS

Hadoop/Big Data: HDFS, Map Reduce, Hive, Pig, YARN, Sqoop, Flume, Oozie, Crunch, Strom, Scala, Kafka, Spark, AWS, R Hadoop

Database/No Sql: SQL.Pl/SQL,HBase,Cassandra

J2EE Frameworks: Hibernate, Spring, JMS, EJB, JSF

XML/Web Services: SOAP/ Rest

Methodologies: Agile, Waterfall

Build Tools: Maven, ANT, Log4j.

Scripting languages: JavaScript, HTML, HTML5, XML, Python, JSP and Servlets

Tools: and other: Rational Rose, Microsoft Visio, CSS

Operating Systems: Linux/Unix, WINDOWS

PROFESSIONAL EXPERIENCE

Confidential, Sunnyvale, CA

Sr. Hadoop Developer

Responsibilities:

  • Involved in Analysis, Design, Development and Testing process based on the new business requirements.
  • Experience managing Cloudera distribution of Hadoop CDH5.5 using Cloudera Manager.
  • Developed java processes for generating different sample log types which generates different random patterns.
  • Developed script for generating flume configuration files for different logs types and for different environments.
  • Worked on setting up flume agents in DEV and UAT environments.
  • Worked on data ingestions, integrating with flume.
  • Experience in refactoring the existing spark batch process for different logs written in Scala.
  • Writing workflows and scheduling using Oozie.
  • Deploying Oozie jobs in DEV and UAT for testing the hourly jobs by parsing the generated sample logs in the same environments.
  • End to end testing the ingestion pipeline in UAT from generating logs to saving the transformed data in HDFS.
  • Responsible in handling streaming data like web server log data using flume.
  • Experience in working with Hive for processing the raw data.
  • Involved in daily SCRUM meetings to discuss the development/progress of Sprints and was active in making scrum meetings more productive.

Environment: Hadoop, Hive, Flume, Scala, Spark, Cloudera, Linux, Maven, Java (JDK1.8), J2EE.

Confidential, Las Angeles, CA

Data Engineer

Responsibilities:

  • Experience working with analysts and product owners for getting exact requirement from clients when they provide some new data sources.
  • Extensively used Amazon s3 as cloud storage for all the ingested and transformed data.
  • Migrated all data and tables from Spark 0.9 to Spark 1.4. Created a parallel process for data all ingestions which uses spark 1.4 and then comparing and validating data to 0.9 there by killing processes in 0.9.
  • Developed java code to get data from different type’s sources which uses REST API, screen scraping from diverse websites depending on type of source and ingested it into a central repository.
  • Responsible for developing Hive queries involving bucketing, partitioning and multiple UDF implementations.
  • Involved in developing UDF’s and UDAF’s to implement customized transformations.
  • Extensively used Sqoop to get data from RDBMS sources like Teradata and Netezza.
  • Validated if data has been normalized and loaded into S3 buckets using Hive queries.
  • Wrote custom Hive UDF’s based on business requirement.
  • Developed an automated system which can import data from any type of source to any type of destination.
  • Developed SparkSQL automation components and responsible for modifying java component to directly connect to thrift server.
  • Experience in writing job flows in python which integrates developed java code and shell scripts which run on Azkaban server.
  • Scheduled and monitored jobs using Azkaban.
  • Experience in spinning up an Amazon EMR cluster with custom configuration and resource allocation.
  • Experience in providing on call support for any cluster issues etc. Also experienced in solving data issues when a job fails in production cluster coz of cluster setup issue, working on it on the fly when analytics team creates an incident.
  • Experience is setting up custom slack notifications where notification can be sent to a group/channel or as a personal message for failure jobs.
  • Developed a POC on comparing CSV files with parquet files on big data sets by performing aggregations and observing time responses.

Environment: Hadoop, Hive, Sqoop, Azkaban, Spark, AWS S3, EMR, Tableau, Linux, Python, Oracle10g, Maven, Java (JDK1.8), J2EE.

Confidential, Portland, OR

Hadoop Developer

Responsibilities:

  • Implemented different aggregations and filtering operations using Java API in Cassandra
  • Responsible for Data modeling in Cassandra and deciding the row key and different column families in Cassandra.
  • Import/export data from Oracle data base to/from HDFS using Sqoop, Hue and JDBC.
  • Gathered data from different sources like Internet, sensors, user behavior using Flume and Kafka and moved to HDFS and implemented Optimized join base using Map Reduce programs.
  • Implemented Custom Input formats dat handle wide range of input files received from java applications to process in Map Reduce.
  • Implemented joins and data aggregation using Apache Crunch.
  • Writing Map Reduce pipeline programs for testing using Apache Crunch.
  • Experienced in unit testing Map Reduce programs using MRUnit.
  • Divided each data set in to corresponding categories by following Map Reduce Binning design pattern.
  • Implemented Filter Mappers to eliminate un-necessary records and perform data and schema validation.
  • Experience in using Pig as an ETL tool for event joins, filters, transformations and pre- aggregations.
  • Created partitions, bucketing across state in Hive to handle structured data.
  • Implemented Dash boards dat handle HiveQL queries internally like Aggregation functions, basic hive operations, and different kind of join operations.
  • Implemented business logic based on state in Hive using Generic UDF's.
  • Involved in creating data-models for customer data using Cassandra Query Language.
  • Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
  • Created production jobs using Oozie work flows dat integrated different actions like Map Reduce, Sqoop, and Hive by utilizing fork and join operations in Oozie.
  • Worked on building BI reports in Tableau with Spark using SparkSQL.
  • Experience in deploying data from various sources into HDFS and building reports using Tableau.
  • Implemented data ingestion and handling clusters in real time processing using Kafka.
  • Experienced in data modelling in hive implementing hive-indexing.
  • Experience in utilizing spark machine learning techniques implemented in Scala.
  • Experienced in configuring maven builds dat integrated dependencies check styles, test coverage's.
  • Responsible for analyzing multi-platform applications using python.
  • Developed Map Reduce jobs in Python for data cleaning and data processing.
  • Designing Test Plans, Test Cases and performed System Testing.

Environment: Hadoop, Map Reduce, Pig, Hive, Sqoop, Oozie, Crunch, Strom, Kafka, Tableau, Cassandra, Linux, Python, R, Oracle10g, Cloudera manager, Maven, MRUnit, JUnit

Confidential, Sylmar, CA

Java/Hadoop Developer

Responsibilities:

  • Involved in Analysis, Design, Development and Testing process of phase 1.
  • Access HBase from Java using Java API to perform CRUD operations on HBase to perform real time analytics.
  • Experienced in loading & transforming large sets of structured, semi-structured and unstructured data to Hadoop.
  • Experienced with different APIs to access HBASE like thrift, Java API.
  • Exported the patterns analyzed back to Teradata using Sqoop.
  • Implemented Region servers to track secure based regions in HBase.
  • Experienced join data from different data sources using Pig Join operations.
  • Experienced with implementing analytics on providing insights using pig sampling.
  • Experienced with handling Sequence file, small file problems, file based structures like log files in Map Reduce.
  • Experienced with Optimizing Hive queries by enabling certain features in Hive.
  • Experienced with implementing incremental updates in hive by fallowing four-step strategies.
  • Experienced with implementing Optimized Hive joins like Map-Side joins and Bucketed Map joins.
  • Integrating bulk data into Cassandra file system using Map Reduce programs.
  • Experienced with moving data from Teradata/ RDBMS systems into HDFS using Sqoop.
  • Integrated Oozie client with java applications to build work flows into application.
  • Implemented Unit test cases using MRUnit, Easy Mock and JUnit.
  • Involved in application design like Sequence Diagrams, Class Diagrams using Microsoft VISIO tool.

Environment: Hadoop, HDFS, Hive, HBase, Map Reduce, Pig, Cassandra Hive, Sqoop, Oozie, HL7, MRUnit, JUnit, UNIX, Shell Scripting, MS Visio

Confidential, Irving, TX

J2EE Developer

Responsibilities:

  • Involved in design of ICD (Interface control design) sequence diagrams, class diagrams and data flow diagrams.
  • Involved in Bug fixes(Refactoring) of written code.
  • Implemented RESTFUL web services with JERSEY implementation for shipping client calls.
  • Worked with Design patterns like Service Factory, Singleton and Factory Pattern.
  • Worked with Servlets transformation layer.
  • Worked with RESTFUL service for providing the services in JSON.
  • Worked with core java 1.6 in process layer.
  • In transaction layer used calling backend transaction EJB object using the service factory pattern.
  • Used Confidential Logger for logging purpose.
  • Worked with JAXB for marshaling and parsing of data from back end transaction.
  • Involved in writing J-Unit test cases.
  • Total project was built with Ant. Involved in writing the build script and build properties.
  • Used SVN for source code versioning and code repository.

Environment: Java (JDK1.6), J2EE, Eclipse, JSP, JavaScript, JSTL, Confidential Logger, XML, JAX-B,EJB,Web Logic 10.3,Restful, Rational Rose, JUnit, Ant, SVN.

Confidential

Java/J2EE Developer

Responsibilities:

  • Involved in all the phases of the life cycle of the project from requirements gathering to quality assurance testing.
  • Developed Class diagrams, Sequence diagrams using Rational Rose.
  • Responsible in developing Rich Web Interface modules with Struts tags,JSP, JSTL, CSS, JavaScript, Ajax, GWT.
  • Developed presentation layer using Struts framework, and performed validations using Struts Validator plugin.
  • Created SQL script for the Oracle database
  • Implemented the Business logic using Java Spring Transaction Spring AOP.
  • Implemented persistence layer using Spring JDBC to store and update data in database.
  • Produced web service using WSDL/SOAP standard.
  • Implemented J2EE design patterns like Singleton Pattern with Factory Pattern.
  • Extensively involved in the creation of the Session Beans and MDB, using EJB 3.0.
  • Used Hibernate framework for Persistence layer.
  • Extensively involved in writing Stored Procedures for data retrieval and data storage and updates in Oracle database using Hibernate.
  • Deployed and built the application using Maven.
  • Performed testing using JUnit.
  • Extensively used Log4j for logging throughout the application.
  • Produced a Web service using REST with Jersey implementation for providing customer information.
  • Used SVN for source code versioning and code repository.

Environment: Java (JDK1.5), J2EE, Eclipse, JSP, JavaScript, JSTL, Ajax, GWT, Log4j, CSS, XML, Spring, EJB, MDB, Hibernate, WebLogic, REST, Rational Rose, JUnit, Maven, JIRA, SVN.

We'd love your feedback!