Sr. Hadoop Developer Resume
Sunnyvale, CA
SUMMARY
- Around 9 years of experience in Application analysis, Design, Development, Maintenance and Supporting web, Client - server based applications in Java/J2EE technologies which includes 3 years of experience wif Big Data and Hadoop related components like HDFS, Map Reduce, Pig, Hive, YARN, Sqoop, Flume, Crunch, Spark, Strom, Scala, Kafka.
- Experience in multiple Hadoop distributions like Cloudera, MapR, and Horton works.
- Excellent understanding of NoSQL databases like HBase, Cassandra and MongoDB.
- Experience on working structured, unstructured data wif various file formats such as Avro data files, xml files, JSON files, sequence files using Map Reduce programs.
- Work experience wif cloud configurations like Amazon web services (AWS).
- Implemented custom business logic and performed join optimization, secondary sorting, custom sorting using Map Reduce programs.
- Experienced testing and running of Map Reduce pipelines on Apache Crunch.
- Expertise in Data ingestion using SQOOP, Apache Kafka, Spark Streaming and FLUME
- Implemented business logic using Pig scripts. Wrote custom Pig UDF’s to analyze data
- Performed different ETL operations using Pig for joining operations and transformations on data to join, clean, aggregate and analyze data.
- Experience wif Oozie Workflow Engine to automate and parallelize Hadoop, Map Reduce and Pig jobs
- Extensive experience wif wiring SQL queries using HiveQL to perform analytics on structured data.
- Experience in performing data validation using HIVE dynamic partitioning and bucketing.
- Experienced in importing and exporting data between RDBMS and Teradata into HDFS using Sqoop.
- Experienced in handling streaming data like web server log data using flume.
- Worked on Cassandra database and related web services for storing unstructured data.
- Good noledge analyzing data using Python development and scripting for Hadoop Streaming.
- Experience in implementing algorithms for analyzing using spark.
- Experience in implementing Spark using Scala and SparkSQL for faster processing of data.
- Experience in getting data from various sources into HDFS and building reports using Tableau.
- Experience in creating tables on top of data on AWS S3 obtained from different data sources and providing them to analytics team building reports using Tableau.
- Extensive Hands on experience wif Accessing and perform CURD operations against HBase data using Java API and implementing time series data management.
- Involved in various data mining tasks such as pattern mining, classification and clustering techniques.
- Expert noledge over J2EE Design Patterns like MVC Architecture, Singleton, Factory Pattern, Front Controller, Session Facade, Business Delegate and Data Access Object for building J2EE Applications.
- Experienced in J2EE, Spring, Hibernate, SOAP/Rest web services,JMS, JNDI, EJB, JAX-WS.
- Expertise wif Application servers and web servers like Oracle WebLogic, IBM WebSphere, Apache Tomcat, JBOSS and VMware.
- Proven expertise in implementing IOC/Dependency Injection features in various aspects of Spring Framework.
- Experienced in developing teh unit test cases using MRUnit, JUnit and Easy Mock.
- Knowledge on Splunk for logging mechanism.
- Knowledge on Build tool Jenkins.
- Experience in using Maven and ANT for build automation.
- Experience in using version control and configuration management tools like SVN, CVS.
- Experience working in environments using Agile (SCRUM) and Waterfall methodologies.
- Experience in designing applications using UML Diagrams like Class Diagram, Component Diagram, Sequence Diagrams, and Deployment Diagram using MS Visio, Rational Rose.
- Expertise in database modeling, administration and development usingSQL and PL/SQL in Oracle (8i, 9i and 10g), MySQL, Teradata, DB2and SQL Server environments.
TECHNICAL SKILLS
Hadoop/Big Data: HDFS, Map Reduce, Hive, Pig, YARN, Sqoop, Flume, Oozie, Crunch, Strom, Scala, Kafka, Spark, AWS, R Hadoop
Database/No Sql: SQL.Pl/SQL,HBase,Cassandra
J2EE Frameworks: Hibernate, Spring, JMS, EJB, JSF
XML/Web Services: SOAP/ Rest
Methodologies: Agile, Waterfall
Build Tools: Maven, ANT, Log4j.
Scripting languages: JavaScript, HTML, HTML5, XML, Python, JSP and Servlets
Tools: and other: Rational Rose, Microsoft Visio, CSS
Operating Systems: Linux/Unix, WINDOWS
PROFESSIONAL EXPERIENCE
Confidential, Sunnyvale, CA
Sr. Hadoop Developer
Responsibilities:
- Involved in Analysis, Design, Development and Testing process based on teh new business requirements.
- Experience managing Cloudera distribution of Hadoop CDH5.5 using Cloudera Manager.
- Developed java processes for generating different sample log types which generates different random patterns.
- Developed script for generating flume configuration files for different logs types and for different environments.
- Worked on setting up flume agents in DEV and UAT environments.
- Worked on data ingestions, integrating wif flume.
- Experience in refactoring teh existing spark batch process for different logs written in Scala.
- Writing workflows and scheduling using Oozie.
- Deploying Oozie jobs in DEV and UAT for testing teh hourly jobs by parsing teh generated sample logs in teh same environments.
- End to end testing teh ingestion pipeline in UAT from generating logs to saving teh transformed data in HDFS.
- Responsible in handling streaming data like web server log data using flume.
- Experience in working wif Hive for processing teh raw data.
- Involved in daily SCRUM meetings to discuss teh development/progress of Sprints and was active in making scrum meetings more productive.
Environment: Hadoop, Hive, Flume, Scala, Spark, Cloudera, Linux, Maven, Java (JDK1.8), J2EE.
Confidential, Las Angeles, CA
Data Engineer
Responsibilities:
- Experience working wif analysts and product owners for getting exact requirement from clients when they provide some new data sources.
- Extensively used Amazon s3 as cloud storage for all teh ingested and transformed data.
- Migrated all data and tables from Spark 0.9 to Spark 1.4. Created a parallel process for data all ingestions which uses spark 1.4 and then comparing and validating data to 0.9 there by killing processes in 0.9.
- Developed java code to get data from different type’s sources which uses REST API, screen scraping from diverse websites depending on type of source and ingested it into a central repository.
- Responsible for developing Hive queries involving bucketing, partitioning and multiple UDF implementations.
- Involved in developing UDF’s and UDAF’s to implement customized transformations.
- Extensively used Sqoop to get data from RDBMS sources like Teradata and Netezza.
- Validated if data TEMPhas been normalized and loaded into S3 buckets using Hive queries.
- Wrote custom Hive UDF’s based on business requirement.
- Developed an automated system which can import data from any type of source to any type of destination.
- Developed SparkSQL automation components and responsible for modifying java component to directly connect to thrift server.
- Experience in writing job flows in python which integrates developed java code and shell scripts which run on Azkaban server.
- Scheduled and monitored jobs using Azkaban.
- Experience in spinning up an Amazon EMR cluster wif custom configuration and resource allocation.
- Experience in providing on call support for any cluster issues etc. Also experienced in solving data issues when a job fails in production cluster because of cluster setup issue, working on it on teh fly when analytics team creates an incident.
- Experience is setting up custom slack notifications where notification can be sent to a group/channel or as a personal message for failure jobs.
- Developed a POC on comparing CSV files wif parquet files on big data sets by performing aggregations and observing time responses.
Environment: Hadoop, Hive, Sqoop, Azkaban, Spark, AWS S3, EMR, Tableau, Linux, Python, Oracle10g, Maven, Java (JDK1.8), J2EE.
Confidential, Portland, OR
Hadoop Developer
Responsibilities:
- Implemented different aggregations and filtering operations using Java API in Cassandra
- Responsible for Data modeling in Cassandra and deciding teh row key and different column families in Cassandra.
- Import/export data from Oracle data base to/from HDFS using Sqoop, Hue and JDBC.
- Gathered data from different sources like Internet, sensors, user behavior using Flume and Kafka and moved to HDFS and implemented Optimized join base using Map Reduce programs.
- Implemented Custom Input formats dat handle wide range of input files received from java applications to process in Map Reduce.
- Implemented joins and data aggregation using Apache Crunch.
- Writing Map Reduce pipeline programs for testing using Apache Crunch.
- Experienced in unit testing Map Reduce programs using MRUnit.
- Divided each data set in to corresponding categories by following Map Reduce Binning design pattern.
- Implemented Filter Mappers to eliminate un-necessary records and perform data and schema validation.
- Experience in using Pig as an ETL tool for event joins, filters, transformations and pre- aggregations.
- Created partitions, bucketing across state in Hive to handle structured data.
- Implemented Dash boards dat handle HiveQL queries internally like Aggregation functions, basic hive operations, and different kind of join operations.
- Implemented business logic based on state in Hive using Generic UDF's.
- Involved in creating data-models for customer data using Cassandra Query Language.
- Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
- Created production jobs using Oozie work flows dat integrated different actions like Map Reduce, Sqoop, and Hive by utilizing fork and join operations in Oozie.
- Worked on building BI reports in Tableau wif Spark using SparkSQL.
- Experience in deploying data from various sources into HDFS and building reports using Tableau.
- Implemented data ingestion and handling clusters in real time processing using Kafka.
- Experienced in data modelling in hive implementing hive-indexing.
- Experience in utilizing spark machine learning techniques implemented in Scala.
- Experienced in configuring maven builds dat integrated dependencies check styles, test coverage's.
- Responsible for analyzing multi-platform applications using python.
- Developed Map Reduce jobs in Python for data cleaning and data processing.
- Designing Test Plans, Test Cases and performed System Testing.
Environment: Hadoop, Map Reduce, Pig, Hive, Sqoop, Oozie, Crunch, Strom, Kafka, Tableau, Cassandra, Linux, Python, R, Oracle10g, Cloudera manager, Maven, MRUnit, JUnit
Confidential, Sylmar, CA
Java/Hadoop Developer
Responsibilities:
- Involved in Analysis, Design, Development and Testing process of phase 1.
- Access HBase from Java using Java API to perform CRUD operations on HBase to perform real time analytics.
- Experienced in loading & transforming large sets of structured, semi-structured and unstructured data to Hadoop.
- Experienced wif different APIs to access HBASE like thrift, Java API.
- Exported teh patterns analyzed back to Teradata using Sqoop.
- Implemented Region servers to track secure based regions in HBase.
- Experienced join data from different data sources using Pig Join operations.
- Experienced wif implementing analytics on providing insights using pig sampling.
- Experienced wif handling Sequence file, small file problems, file based structures like log files in Map Reduce.
- Experienced wif Optimizing Hive queries by enabling certain features in Hive.
- Experienced wif implementing incremental updates in hive by fallowing four-step strategies.
- Experienced wif implementing Optimized Hive joins like Map-Side joins and Bucketed Map joins.
- Integrating bulk data into Cassandra file system using Map Reduce programs.
- Experienced wif moving data from Teradata/ RDBMS systems into HDFS using Sqoop.
- Integrated Oozie client wif java applications to build work flows into application.
- Implemented Unit test cases using MRUnit, Easy Mock and JUnit.
- Involved in application design like Sequence Diagrams, Class Diagrams using Microsoft VISIO tool.
Environment: Hadoop, HDFS, Hive, HBase, Map Reduce, Pig, Cassandra Hive, Sqoop, Oozie, HL7, MRUnit, JUnit, UNIX, Shell Scripting, MS Visio
Confidential, Irving, TX
J2EE Developer
Responsibilities:
- Involved in design of ICD (Interface control design) sequence diagrams, class diagrams and data flow diagrams.
- Involved in Bug fixes(Refactoring) of written code.
- Implemented RESTFUL web services wif JERSEY implementation for shipping client calls.
- Worked wif Design patterns like Service Factory, Singleton and Factory Pattern.
- Worked wif Servlets transformation layer.
- Worked wif RESTFUL service for providing teh services in JSON.
- Worked wif core java 1.6 in process layer.
- In transaction layer used calling backend transaction EJB object using teh service factory pattern.
- Used Confidential Logger for logging purpose.
- Worked wif JAXB for marshaling and parsing of data from back end transaction.
- Involved in writing J-Unit test cases.
- Total project was built wif Ant. Involved in writing teh build script and build properties.
- Used SVN for source code versioning and code repository.
Environment: Java (JDK1.6), J2EE, Eclipse, JSP, JavaScript, JSTL, Confidential Logger, XML, JAX-B,EJB,Web Logic 10.3,Restful, Rational Rose, JUnit, Ant, SVN.
Confidential
Java/J2EE Developer
Responsibilities:
- Involved in all teh phases of teh life cycle of teh project from requirements gathering to quality assurance testing.
- Developed Class diagrams, Sequence diagrams using Rational Rose.
- Responsible in developing Rich Web Interface modules wif Struts tags,JSP, JSTL, CSS, JavaScript, Ajax, GWT.
- Developed presentation layer using Struts framework, and performed validations using Struts Validator plugin.
- Created SQL script for teh Oracle database
- Implemented teh Business logic using Java Spring Transaction Spring AOP.
- Implemented persistence layer using Spring JDBC to store and update data in database.
- Produced web service using WSDL/SOAP standard.
- Implemented J2EE design patterns like Singleton Pattern wif Factory Pattern.
- Extensively involved in teh creation of teh Session Beans and MDB, using EJB 3.0.
- Used Hibernate framework for Persistence layer.
- Extensively involved in writing Stored Procedures for data retrieval and data storage and updates in Oracle database using Hibernate.
- Deployed and built teh application using Maven.
- Performed testing using JUnit.
- Extensively used Log4j for logging throughout teh application.
- Produced a Web service using REST wif Jersey implementation for providing customer information.
- Used SVN for source code versioning and code repository.
Environment: Java (JDK1.5), J2EE, Eclipse, JSP, JavaScript, JSTL, Ajax, GWT, Log4j, CSS, XML, Spring, EJB, MDB, Hibernate, WebLogic, REST, Rational Rose, JUnit, Maven, JIRA, SVN.
