We provide IT Staff Augmentation Services!

Spark/hadoop Developer Resume

5.00/5 (Submit Your Rating)

NyC

SUMMARY:

  • Extensive IT experience of over 7 years with multinational clients which includes 4 years of Hadoop related architecture experience developing Big data / Hadoop applications.
  • Hands on experience with the Hadoop stack (MapReduce, Pig, Hive, Sqoop, HBase, Flume, Oozie).
  • Proven Expertise in performing analytics on Big Data using Map Reduce, Hive and Pig.
  • Implemented POC to migrate map reduce jobs into Spark RDD transformations using Scala
  • Developed Apache Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying.
  • Experienced with performing real time analytics on NoSQL data bases like HBase, MongoDB and Cassandra.
  • Proficient knowledge in working with Impala, Storm and Kafka.
  • Experienced with Dimensional modeling, Data migration, Data cleansing, Data profiling, and ETL Processes features for data warehouses.
  • Worked with Oozie work flow engine to schedule time based jobs to perform multiple actions.
  • Experienced in importing and exporting data from RDBMS into HDFS using Sqoop.
  • Analyzed large amounts of data sets using Pig scripts and Hive scripts.
  • Logical Implementation and interaction with HBase.
  • Experienced in writing MapReduce programs and UDFs for both Hive and Pig in Java.
  • Used Flume to channel data from different sources to HDFS.
  • Experience with configuration of Hadoop Ecosystem components: Hive, HBase, Pig, Sqoop, Mahout and Flume.
  • Good experience in Hive partitioning, bucketing and perform different types of joins on Hive tables and implementing Hive SerDe like JSON and Avro.
  • Experience in Performance Tuning, Optimization and Customization.
  • Supported MapReduce Programs running on the cluster and wrote custom MapReduce Scripts for Data Processing in Java.
  • Good knowledge in Apache Crunch and Hadoop HDFS Admin Shell commands.
  • Experience with Unit Testing Map Reduce programs using MRUnit, JUnit and Easy Mock.
  • Experience in Active Development as well as onsite coordination activities in web based, client/server and distributed architecture using Java, J2EE which includes Web services, Spring, Struts 1.0, Hibernate and JSP/Servlets along with incorporating MVC architecture.
  • As part of my assignments, I have been working on projects for leading clients, which includes Understanding client Requirements, Estimations, Analysis of Functional specifications, Technical Design Specification, Review of Technical Design, Development, Testing and Implementation activities.
  • During this period I have also acquired strong knowledge of Software Quality Processes and SDLC (Software Development Life Cycle).
  • Good working knowledge on servers like Tomcat and Websphere.
  • Extensively worked on Java development tools, which includes Eclipse, WSAD and JBuilder.
  • I am Innovative and self - motivated with strong communication and interpersonal skills, highly adaptable and highly customer focused.
  • Ability to work in teams as well as an individual, quick learner and able to meet deadlines.

TECHNICAL SKILLS:

Hadoop Ecosystem: Hadoop, MapReduce, YARN, Spark, Sqoop, Hive, Oozie, PIG, HDFS, Flume, Impala, Storm, Kafka

Programming Languages: C, C++, JAVA, Scala SQL, PL/SQL, PIG Latin, HiveQL, Unix shell scripting

Java & J2EE Technologies: Core Java, Servlets, JSP, JDBC, EJB s

Frameworks: Spring, Hibernate, Struts 1/2, EJB, JMS, JUnit, MRUnit

No SQL Databases: HBase, Cassandra and MongoDbDatabases: Oracle 11g/10g/9i, My SQL, DB2, MS SQL Server

Application Server: Apache Tomcat, JBoss, IBM Web sphere, Web Logic

Web Services: RESTful, SOAP, Apache CXF, Apache Axis

Methodologies: Scrum, Agile, Waterfall

WORK EXPERIENCE:

Confidential, NYC

Spark/Hadoop Developer

Roles and responsibilities:

  • Preparing Design Documents (Request-Response Mapping Documents, Hive Mapping Documents).
  • Involved in design Cassandra data model, used CQL (Cassandra Query Language) to perform CRUD operations on Cassandra file system
  • Experienced with batch processing of data sources using Apache Spark and Elastic search.
  • Experienced in implementing Spark RDD transformations, actions to implement business analysis
  • Migrated Hive QL queries on structured into Spark QL to improve performance
  • D eveloped code base to stream data from sample Data files Kafka Kafka Spout Storm Bolt HDFS BOLT
  • D ocumented the data flow form Application Kafka Storm HDFS Hive tables
  • Configured, deployed and maintained a single node storm cluster in DEV environment
  • Developing predictive analytic using Apache Spark Scala APIs
  • Developed solutions to pre-process large sets of structured, semi-structured data, with different file formats (Text file, Avro data files, Sequence files, Xml and JSon files, ORC and Parquet).
  • Handled importing of data from RDBMS into HDFS using Sqoop.
  • Experienced in data cleansing processing using Pig latin operations and UDFs.
  • Experienced in writing Hive Scripts for analyzing data in Hive warehouse using Hive Query Language (HQL).
  • Involved in creating Hive tables, loading with data and writing hive queries to process the data.
  • Created scripts to automate the process of Data Ingestion.
  • Developed PIG scripts for source data validation and transformation.
  • Installed Oozie workflow engine to run multiple Hive and Pig jobs which run independently with time and data availability for analyzing HDFS audit data.
  • Preparing korn Shell jobs and pushing the code to DEV, UAT, PROD environments.
  • Experience in using Testing Frameworks of BigData world, MRUnit, PIGUnit for testing raw data and executed performance scripts.

Tools: and technologies used: HDFS, Apache Spark, Kafka, Cassandra, Storm Hive, Pig, Scala, Java, SqoopSQL, Shell scripting.

Confidential, IL

Sr. Hadoop Developer

Roles and responsibilities:

  • Required Analysis along with business and peers.
  • Preparing Design Documents (Request-Response Mapping Documents, Hive Mapping Documents).
  • Used Flume and Sqoop to load data from multiple sources into HDFS .
  • Handled importing of data from various data sources, performed transformations using Pig and Hive to load data into HDFS.
  • Experience in joining raw data with the reference data using Pig scripting and Hive scripting.
  • Created Oozie workflow engine to run multiple Hive and Pig jobs.
  • Created O0zie coordinated workflow to execute Sqoop incremental job daily.
  • Coordinated with ETL engineers for ingestion of various data sources.
  • Analyzed the data by performing data profiling.
  • Performing the testing in various environments and providing reports to the business

Tools: and technologies used: Hadoop, HDFS, Hive, Pig, Sqoop, Flume, Oozie.

Confidential, Madison, WI

Sr. Hadoop Developer

Roles and responsibilities:

  • Involved in Installing, Configuring Hadoop Eco System, Cloudera Manager using CDH4 Distribution.
  • Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
  • Processed Multiple Data sources input to same Reducer using Generic Writable and Multi Input format.
  • Created Data Pipeline of Map Reduce programs using Chained Mappers.
  • Visualize the HDFS data to customer using BI tool with the help of Hive ODBC Driver.
  • Familiarity with a NoSQL database such as MongoDb.
  • Implemented Optimized join base by joining different data sets to get top claims based on state using Map Reduce.
  • Worked Big data processing of clinical and non clinical data using Map Reduce.
  • Implemented complex map reduce programs to perform joins on the Map side using Distributed Cache in Java.
  • Responsible for importing log files from various sources into HDFS using Flume.
  • Created customized BI tool for manager team that perform Query analytics using HiveQL.
  • Used Hive and Pig to generate BI reports.
  • Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
  • Created Partitions, Buckets based on State to further process using Bucket based Hive joins.
  • Created Hive Generic UDF's to process business logic that varies based on policy.
  • Moved Relational Data base data using Sqoop into Hive Dynamic partition tables using staging tables.
  • Worked on custom Pig Loaders and storage classes to work with variety of data formats such as JSON and XML file formats.
  • Experienced with different kind of compression techniques like LZO, GZip, Snappy.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
  • Developed Unit test cases using Junit, Easy Mock and MRUnit testing frameworks.
  • Experienced in Monitoring Cluster using Cloudera manager.

Tools: and technologies used: Hadoop, HDFS, HBase, MongoDb, MapReduce, Java, Hive, Pig, Sqoop, Flume, Oozie, Hue, SQL, ETL, Cloudera Manager, MySQL.

Confidential, Austin, TX

Hadoop Developer

Roles and responsibilities:

  • Worked on importing data from various sources and performed transformations using MapReduce, Hive to load data into HDFS.
  • Configured Sqoop jobs to import data from RDBMS into HDFS using Oozie workflows.
  • Worked on setting up Pig, Hive and HBase on multiple nodes and developed using Pig, Hive, HBase and MapReduce.
  • Solved small file problem using Sequence files processing in Map Reduce.
  • Written various Hive and Pig scripts.
  • Created HBase tables to store variable data formats coming from different portfolios.
  • Performed real time analytics on HBase using Java API and Rest API.
  • Implemented HBase Co-processors to notify Support team when inserting data into HBase Tables.
  • Worked on compression mechanisms to optimize MapReduce Jobs.
  • Analyzed the customer behavior by performing click stream analysis and to ingest the data used flume.
  • Experienced with working on Avro Data files using Avro Serialization system.
  • Implemented business logic by writing UDF's in Java and used various UDF's from Piggybanks and other sources.
  • Continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.

Tools: and technologies used: Horton works, Map Reduce, HBase, HDFS, Hive, Pig, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Java (jdk 1.6), Eclipse

Confidential, Champaign, IL

Big Data Analyst/ Java Developer

Roles and responsibilities:

  • Installed and configured Apache Hadoop to test the maintenance of log files in Hadoop cluster.
  • Installed and configured Hive, Pig, Sqoop, and Oozie on the Hadoop cluster.
  • Installed Oozie Workflow engine to run multiple Hive and Pig Jobs.
  • Developed multiple MapReduce jobs in Java for data cleansing and preprocessing.
  • Developed Simple to complex Map/Reduce Jobs using Hive and Pig.
  • Involved in loading data from UNIX file system to HDFS.
  • Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Provided quick response to ad hoc internal and external client requests for data and experienced in creating ad hoc reports.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Migration of ETL processes from Oracle to Hive to test the easy data manipulation.
  • Performed optimization on Pig scripts and Hive queries increase efficiency and add new features to existing code.
  • Stored and retrieved data from data-warehouses using Amazon Redshift.
  • Developed PIG Latin scripts for the analysis of semi structured data.
  • Used Hive and created Hive tables and involved in data loading and writing Hive UDFs.
  • Used Sqoop to import data into HDFS and Hive from other data systems.
  • Installed Oozie workflow engine to run multiple Hive.
  • Generated aggregations and groups and visualizations using Tableau.
  • Continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
  • Conducted some unit testing for the development team within the sandbox environment.
  • Developed Hive queries to process the data

Tools: and technologies used: Apache Hadoop, Cloudera Manager, CDH2, CDH3 CentOS, Java, MapReduce, Apache Hama, Eclipse Indigo, Pig, Hive, Sqoop, Oozie and SQL, Struts, JUnit.

Confidential

Software Developer

Roles and responsibilities:

  • Involved in the complete SDLC software development life cycle of the application from requirement gathering and analysis to testing and maintenance.
  • Implemented the User Login logic using Spring MVC framework encouraging application architectures based on the Model View Controller design paradigm.
  • Generated Hibernate Mapping files and created the data model using mapping files.
  • Developed UI using JavaScript, JSP, HTML and CSS for interactive cross browser functionality and complex user interface.
  • Used Struts Tiles and Validator framework in developing the applications.
  • Developed action classes and form beans and configured the struts-config.xml
  • Provided client side validations using Struts Validator framework and JavaScript
  • Created business logic using servlets and session beans and deployed them on Apache Tomcat server.
  • Created complex SQL Queries, PL/SQL Stored procedures and functions for back end.
  • Prepared the functional, design and test case specifications.
  • Performed unit testing, system testing and integration testing.
  • Developed unit test cases. Used JUnit for unit testing of the application.
  • Provided Technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects. Resolved more priority defects as per the schedule.

Tools: and technologies used: Java, Spring MVC, Struts, Hibernate, JSP, Servlets, WebServices, Apache Tomcat, Oracle, JUnit, SQL.

Confidential

Software Trainee

Roles and responsibilities:

  • Involved in the complete development, testing and maintenance process of the application.
  • Responsible for gathering the requirements doing the analysis and formulating the requirements specifications with the consistent inputs/requirements.
  • Developed JSP as an application controller.
  • Designed and developed HTML front end screens and validated forms using JavaScript.
  • Used Frames and Cascading Style Sheets (CSS) to give a better view to the Web Pages.
  • Deployed the web application on Web Logic server.
  • Used JDBC for database connectivity.
  • Developed necessary SQL queries for database transactions.
  • Involved in testing, implementation and documentation.
  • Written Java script code for Input Validation.
  • Front End was built using JSPs, JavaScript and HTML.
  • Built Custom Tags for JSPs.
  • Built the report module on reports based from Crystal reports.
  • Integrating data from multiple data sources.
  • Generating schema difference reports for database using toad.

Tools: and technologies used: Java, JSP, Web Logic 5.1, HTML, JavaScript, JDBC and SQL, PL/SQL, UNIX.

We'd love your feedback!