Spark/hadoop Developer Resume
NyC
SUMMARY:
- Extensive IT experience of over 7 years with multinational clients which includes 4 years of Hadoop related architecture experience developing Big data / Hadoop applications.
- Hands on experience with the Hadoop stack (MapReduce, Pig, Hive, Sqoop, HBase, Flume, Oozie).
- Proven Expertise in performing analytics on Big Data using Map Reduce, Hive and Pig.
- Implemented POC to migrate map reduce jobs into Spark RDD transformations using Scala
- Developed Apache Spark jobs using Scala in test environment for faster data processing and used Spark SQL for querying.
- Experienced with performing real time analytics on NoSQL data bases like HBase, MongoDB and Cassandra.
- Proficient knowledge in working with Impala, Storm and Kafka.
- Experienced with Dimensional modeling, Data migration, Data cleansing, Data profiling, and ETL Processes features for data warehouses.
- Worked with Oozie work flow engine to schedule time based jobs to perform multiple actions.
- Experienced in importing and exporting data from RDBMS into HDFS using Sqoop.
- Analyzed large amounts of data sets using Pig scripts and Hive scripts.
- Logical Implementation and interaction with HBase.
- Experienced in writing MapReduce programs and UDFs for both Hive and Pig in Java.
- Used Flume to channel data from different sources to HDFS.
- Experience with configuration of Hadoop Ecosystem components: Hive, HBase, Pig, Sqoop, Mahout and Flume.
- Good experience in Hive partitioning, bucketing and perform different types of joins on Hive tables and implementing Hive SerDe like JSON and Avro.
- Experience in Performance Tuning, Optimization and Customization.
- Supported MapReduce Programs running on the cluster and wrote custom MapReduce Scripts for Data Processing in Java.
- Good knowledge in Apache Crunch and Hadoop HDFS Admin Shell commands.
- Experience with Unit Testing Map Reduce programs using MRUnit, JUnit and Easy Mock.
- Experience in Active Development as well as onsite coordination activities in web based, client/server and distributed architecture using Java, J2EE which includes Web services, Spring, Struts 1.0, Hibernate and JSP/Servlets along with incorporating MVC architecture.
- As part of my assignments, I have been working on projects for leading clients, which includes Understanding client Requirements, Estimations, Analysis of Functional specifications, Technical Design Specification, Review of Technical Design, Development, Testing and Implementation activities.
- During this period I have also acquired strong knowledge of Software Quality Processes and SDLC (Software Development Life Cycle).
- Good working knowledge on servers like Tomcat and Websphere.
- Extensively worked on Java development tools, which includes Eclipse, WSAD and JBuilder.
- I am Innovative and self - motivated with strong communication and interpersonal skills, highly adaptable and highly customer focused.
- Ability to work in teams as well as an individual, quick learner and able to meet deadlines.
TECHNICAL SKILLS:
Hadoop Ecosystem: Hadoop, MapReduce, YARN, Spark, Sqoop, Hive, Oozie, PIG, HDFS, Flume, Impala, Storm, Kafka
Programming Languages: C, C++, JAVA, Scala SQL, PL/SQL, PIG Latin, HiveQL, Unix shell scripting
Java & J2EE Technologies: Core Java, Servlets, JSP, JDBC, EJB s
Frameworks: Spring, Hibernate, Struts 1/2, EJB, JMS, JUnit, MRUnit
No SQL Databases: HBase, Cassandra and MongoDbDatabases: Oracle 11g/10g/9i, My SQL, DB2, MS SQL Server
Application Server: Apache Tomcat, JBoss, IBM Web sphere, Web Logic
Web Services: RESTful, SOAP, Apache CXF, Apache Axis
Methodologies: Scrum, Agile, Waterfall
WORK EXPERIENCE:
Confidential, NYC
Spark/Hadoop Developer
Roles and responsibilities:
- Preparing Design Documents (Request-Response Mapping Documents, Hive Mapping Documents).
- Involved in design Cassandra data model, used CQL (Cassandra Query Language) to perform CRUD operations on Cassandra file system
- Experienced with batch processing of data sources using Apache Spark and Elastic search.
- Experienced in implementing Spark RDD transformations, actions to implement business analysis
- Migrated Hive QL queries on structured into Spark QL to improve performance
- D eveloped code base to stream data from sample Data files Kafka Kafka Spout Storm Bolt HDFS BOLT
- D ocumented the data flow form Application Kafka Storm HDFS Hive tables
- Configured, deployed and maintained a single node storm cluster in DEV environment
- Developing predictive analytic using Apache Spark Scala APIs
- Developed solutions to pre-process large sets of structured, semi-structured data, with different file formats (Text file, Avro data files, Sequence files, Xml and JSon files, ORC and Parquet).
- Handled importing of data from RDBMS into HDFS using Sqoop.
- Experienced in data cleansing processing using Pig latin operations and UDFs.
- Experienced in writing Hive Scripts for analyzing data in Hive warehouse using Hive Query Language (HQL).
- Involved in creating Hive tables, loading with data and writing hive queries to process the data.
- Created scripts to automate the process of Data Ingestion.
- Developed PIG scripts for source data validation and transformation.
- Installed Oozie workflow engine to run multiple Hive and Pig jobs which run independently with time and data availability for analyzing HDFS audit data.
- Preparing korn Shell jobs and pushing the code to DEV, UAT, PROD environments.
- Experience in using Testing Frameworks of BigData world, MRUnit, PIGUnit for testing raw data and executed performance scripts.
Tools: and technologies used: HDFS, Apache Spark, Kafka, Cassandra, Storm Hive, Pig, Scala, Java, SqoopSQL, Shell scripting.
Confidential, IL
Sr. Hadoop Developer
Roles and responsibilities:
- Required Analysis along with business and peers.
- Preparing Design Documents (Request-Response Mapping Documents, Hive Mapping Documents).
- Used Flume and Sqoop to load data from multiple sources into HDFS .
- Handled importing of data from various data sources, performed transformations using Pig and Hive to load data into HDFS.
- Experience in joining raw data with the reference data using Pig scripting and Hive scripting.
- Created Oozie workflow engine to run multiple Hive and Pig jobs.
- Created O0zie coordinated workflow to execute Sqoop incremental job daily.
- Coordinated with ETL engineers for ingestion of various data sources.
- Analyzed the data by performing data profiling.
- Performing the testing in various environments and providing reports to the business
Tools: and technologies used: Hadoop, HDFS, Hive, Pig, Sqoop, Flume, Oozie.
Confidential, Madison, WI
Sr. Hadoop Developer
Roles and responsibilities:
- Involved in Installing, Configuring Hadoop Eco System, Cloudera Manager using CDH4 Distribution.
- Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
- Processed Multiple Data sources input to same Reducer using Generic Writable and Multi Input format.
- Created Data Pipeline of Map Reduce programs using Chained Mappers.
- Visualize the HDFS data to customer using BI tool with the help of Hive ODBC Driver.
- Familiarity with a NoSQL database such as MongoDb.
- Implemented Optimized join base by joining different data sets to get top claims based on state using Map Reduce.
- Worked Big data processing of clinical and non clinical data using Map Reduce.
- Implemented complex map reduce programs to perform joins on the Map side using Distributed Cache in Java.
- Responsible for importing log files from various sources into HDFS using Flume.
- Created customized BI tool for manager team that perform Query analytics using HiveQL.
- Used Hive and Pig to generate BI reports.
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
- Created Partitions, Buckets based on State to further process using Bucket based Hive joins.
- Created Hive Generic UDF's to process business logic that varies based on policy.
- Moved Relational Data base data using Sqoop into Hive Dynamic partition tables using staging tables.
- Worked on custom Pig Loaders and storage classes to work with variety of data formats such as JSON and XML file formats.
- Experienced with different kind of compression techniques like LZO, GZip, Snappy.
- Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java map-reduce Hive, Pig, and Sqoop.
- Developed Unit test cases using Junit, Easy Mock and MRUnit testing frameworks.
- Experienced in Monitoring Cluster using Cloudera manager.
Tools: and technologies used: Hadoop, HDFS, HBase, MongoDb, MapReduce, Java, Hive, Pig, Sqoop, Flume, Oozie, Hue, SQL, ETL, Cloudera Manager, MySQL.
Confidential, Austin, TX
Hadoop Developer
Roles and responsibilities:
- Worked on importing data from various sources and performed transformations using MapReduce, Hive to load data into HDFS.
- Configured Sqoop jobs to import data from RDBMS into HDFS using Oozie workflows.
- Worked on setting up Pig, Hive and HBase on multiple nodes and developed using Pig, Hive, HBase and MapReduce.
- Solved small file problem using Sequence files processing in Map Reduce.
- Written various Hive and Pig scripts.
- Created HBase tables to store variable data formats coming from different portfolios.
- Performed real time analytics on HBase using Java API and Rest API.
- Implemented HBase Co-processors to notify Support team when inserting data into HBase Tables.
- Worked on compression mechanisms to optimize MapReduce Jobs.
- Analyzed the customer behavior by performing click stream analysis and to ingest the data used flume.
- Experienced with working on Avro Data files using Avro Serialization system.
- Implemented business logic by writing UDF's in Java and used various UDF's from Piggybanks and other sources.
- Continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
Tools: and technologies used: Horton works, Map Reduce, HBase, HDFS, Hive, Pig, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Java (jdk 1.6), Eclipse
Confidential, Champaign, IL
Big Data Analyst/ Java Developer
Roles and responsibilities:
- Installed and configured Apache Hadoop to test the maintenance of log files in Hadoop cluster.
- Installed and configured Hive, Pig, Sqoop, and Oozie on the Hadoop cluster.
- Installed Oozie Workflow engine to run multiple Hive and Pig Jobs.
- Developed multiple MapReduce jobs in Java for data cleansing and preprocessing.
- Developed Simple to complex Map/Reduce Jobs using Hive and Pig.
- Involved in loading data from UNIX file system to HDFS.
- Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Provided quick response to ad hoc internal and external client requests for data and experienced in creating ad hoc reports.
- Responsible for building scalable distributed data solutions using Hadoop.
- Migration of ETL processes from Oracle to Hive to test the easy data manipulation.
- Performed optimization on Pig scripts and Hive queries increase efficiency and add new features to existing code.
- Stored and retrieved data from data-warehouses using Amazon Redshift.
- Developed PIG Latin scripts for the analysis of semi structured data.
- Used Hive and created Hive tables and involved in data loading and writing Hive UDFs.
- Used Sqoop to import data into HDFS and Hive from other data systems.
- Installed Oozie workflow engine to run multiple Hive.
- Generated aggregations and groups and visualizations using Tableau.
- Continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
- Conducted some unit testing for the development team within the sandbox environment.
- Developed Hive queries to process the data
Tools: and technologies used: Apache Hadoop, Cloudera Manager, CDH2, CDH3 CentOS, Java, MapReduce, Apache Hama, Eclipse Indigo, Pig, Hive, Sqoop, Oozie and SQL, Struts, JUnit.
Confidential
Software Developer
Roles and responsibilities:
- Involved in the complete SDLC software development life cycle of the application from requirement gathering and analysis to testing and maintenance.
- Implemented the User Login logic using Spring MVC framework encouraging application architectures based on the Model View Controller design paradigm.
- Generated Hibernate Mapping files and created the data model using mapping files.
- Developed UI using JavaScript, JSP, HTML and CSS for interactive cross browser functionality and complex user interface.
- Used Struts Tiles and Validator framework in developing the applications.
- Developed action classes and form beans and configured the struts-config.xml
- Provided client side validations using Struts Validator framework and JavaScript
- Created business logic using servlets and session beans and deployed them on Apache Tomcat server.
- Created complex SQL Queries, PL/SQL Stored procedures and functions for back end.
- Prepared the functional, design and test case specifications.
- Performed unit testing, system testing and integration testing.
- Developed unit test cases. Used JUnit for unit testing of the application.
- Provided Technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects. Resolved more priority defects as per the schedule.
Tools: and technologies used: Java, Spring MVC, Struts, Hibernate, JSP, Servlets, WebServices, Apache Tomcat, Oracle, JUnit, SQL.
Confidential
Software Trainee
Roles and responsibilities:
- Involved in the complete development, testing and maintenance process of the application.
- Responsible for gathering the requirements doing the analysis and formulating the requirements specifications with the consistent inputs/requirements.
- Developed JSP as an application controller.
- Designed and developed HTML front end screens and validated forms using JavaScript.
- Used Frames and Cascading Style Sheets (CSS) to give a better view to the Web Pages.
- Deployed the web application on Web Logic server.
- Used JDBC for database connectivity.
- Developed necessary SQL queries for database transactions.
- Involved in testing, implementation and documentation.
- Written Java script code for Input Validation.
- Front End was built using JSPs, JavaScript and HTML.
- Built Custom Tags for JSPs.
- Built the report module on reports based from Crystal reports.
- Integrating data from multiple data sources.
- Generating schema difference reports for database using toad.
Tools: and technologies used: Java, JSP, Web Logic 5.1, HTML, JavaScript, JDBC and SQL, PL/SQL, UNIX.
