We provide IT Staff Augmentation Services!

Hadoop Developer Resume

4.00/5 (Submit Your Rating)

Waltham, MA

SUMMARY

  • Around 7+ years of professional IT experience which includes experience in PL/SQL and Big - Data related technologies in Finance, Healthcare and Retail Industries.
  • Hands on experience on major components in Hadoop Ecosystem like Hadoop HDFS, MapReduce, YARN, Cassandra, IMPALA, Hive, Pig,Datafu, Drill, HBase, Sqoop, Oozie, Flume.
  • Good Knowledge on Map Reduce design patterns.
  • Experience with distributed systems, large-scale non-relational data stores, NoSQL map-reduce systems, data modeling, database performance tuning, and multi-terabyte data warehouses.
  • Responsible for setting up processes for Hadoop based application design and implementation.
  • Extensively worked on Hive, Pig,Datafu,Impala for performing data analysis.
  • Experience in managing HBase database and using it to update/modify the data quickly.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice-versa.
  • Experience in ingesting log data into HDFS using Flume.
  • Experience in managing and reviewing Hadoop log files.
  • Experience with Cloudera, Horton works and MapR distribution.
  • Handling data in various file formats such as Sequential, AVRO, RC, Parquet and ORC.
  • Very good experience in complete project life cycle (design, development, testing and implementation) of Client Server and Web applications.
  • Experience in Object Oriented Analysis, Design (OOAD) and development of software using UML Methodology, good knowledge of J2EE design patterns, Core Java design patterns and MVC design patterns.
  • Involved in developing complex ETL transformation & performance tuning.
  • Good knowledge in MongoDB concepts and its architecture.
  • Hands on experience in application development using Java, RDBMS, and Linux shell scripting.
  • Worked with application team via scrum to provide operational support, install Hadoop updates, patches and version upgrades as required.
  • Excellent working knowledge of System Development Life Cycle (SDLC) and Software Testing Life Cycle (STLC) and Defect Life Cycle..
  • Well versed in installation, configuration, supporting and managing of Big Data and underlying infrastructure of Hadoop Cluster
  • Ability to blend technical expertise with strong Conceptual, Business and Analytical skills to provide quality solutions and result-oriented problem solving technique and leadership skills.

TECHNICAL SKILLS

Operating Systems: LINUX(Ubuntu, CentOS, Redhat), Windows

Big Data/Hadoop: HDFS, Hadoop MapReduce, Hive, Pig,Drill, Sqoop, HBase, Flume,Impala, Oozie,Zookeeper.

Languages: C, Java, Python, SQL/PLSQL, Shell Scripting, HiveQL, Pig Latin, SQL

Methodologies: Agile, Waterfall model

Databases: HBase, CASSANDRA, Mongodb,MySQL, DB2, Oracle 10g, Teradata

Web Tools/Frameworks: HTML, Java Script, XML, ODBC, JDBC, JSP, Servlets

IDE: Eclipse

Application Servers: Apache Tomcat server, Apache HTTP webserver

PROFESSIONAL EXPERIENCE

Confidential, Waltham, MA

Hadoop Developer

Responsibilities:

  • Developed the UDF's to preprocess the data for analysis.
  • Developed workflow of loading data into HDFS from DB2 using Sqoop,Hive for analysis and preprocessing with PIG and Pig Datafu libraries .
  • Developed workflows using custom MapReduce, Pig and Hive for data cleaning.
  • Built reusable Hive UDF libraries for business requirements which enabled users to use these UDF's in Hive querying.
  • Involved in analysis, design, testing phases and responsible for documenting technical specifications.
  • The logs and semi structured content that are stored on HDFS were preprocessed using PIG and the processed data is imported into Hive warehouse which enabled business analysts to write Hive queries.
  • Schedule these jobs with workflow engine like Oozie. Actions can be performed both sequentially and parallel using Oozie.
  • Built wrapper shell scripts to hold this Oozie workflow.
  • Involved in creating Hadoop streaming jobs using Python and Spark.
  • Developed POC for Apache Kafka.
  • Implemented advanced analytical methodologies using Impala.
  • Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing the data onto HDFS
  • Configured big data workflows to run on the top of Hadoop and these workflows comprises of heterogeneous jobs like Pig, Hive, Sqoop and MapReduce.
  • Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using testing library.
  • Worked on NoSQL databases including HBase, Cassandra and MongoDB.
  • Designed a data warehouse using Hive.
  • Collected the logs data from web servers and integrated in to HDFS using Flume.
  • Wrote shell scripts for rolling day-to-day processes and it is automated.
  • Involved in story-driven agile development methodology and actively participated in daily scrum meetings

Environment: Hadoop 2x, HDFS, Tableau, DB2, Map Reduce, Sqoop,Oozie,Cloudera, HBase, Shell Scripting, PIG, HIVE, Impala,Kafka, Core Java,NoSQL, Oracle 11g, PL/SQL, SQL*PLUS, Windows NT, LINUX

Confidential, KS

Hadoop Developer

Responsibilities:

  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Experience in pig scripting and writing pig programs for map reduce.
  • Experienced in defining job flows.
  • Experienced in managing and reviewing Hadooplog files.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Responsible to manage data coming from different sources.
  • Supported Map Reduce Programs those are running on the cluster.
  • Involved in loading data from UNIX file system to HDFS.
  • Installed and configured Hive and also written Hive UDFs.
  • Involved in creating Hive tables, loading with data and writing hive queries that will run internally in map reduce way.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Extensively working on Core Java for MapReduce Jobs.
  • Implemented best income logic using Pig scripts.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
  • Involved in templates and screens in HTML and JavaScript.

Environment: Hadoop, MapReduce, HDFS, Hive, Java, SQL, PIG, Hbase, Sqoop, HTML, XML

Confidential, Kansas City, MO

Hadoop Developer

Responsibilities:

  • Involved in review of functional and non-functional requirements.
  • Installed and configured Hadoop MapReduce, HDFS, Developed multiple MapReduce jobs in java for data cleaning and preprocessing.
  • Installed and configured Pig and also written Pig Latin scripts.
  • Wrote MapReduce job using Pig Latin.
  • Involved in ETL, Data Integration and Migration
  • Imported data using Sqoop to load data from Oracle to HDFS on regular basis.
  • Developing Scripts and Batch Job to schedule various Hadoop Program.
  • Written Hive queries for data analysis to meet the business requirements.
  • Creating Hive tables and working on them using Hive QL.
  • Importing and exporting data into HDFS from Oracle Database and vice versa using Sqoop.
  • Experienced indefining jobflows.
  • Got good experience with NOSQL database HBase, MongoDB.
  • Involved in creating Hive tables, loading the data and writing hive queries that will run internally in a map reduce way.
  • Developed a custom FileSystem plugin for Hadoop so it can access files on Data Platform.
  • The custom File System plugin allows Hadoop MapReduce programs, HBase, Pig and Hive to work unmodified and access files directly.
  • Designed and implemented MapReduce-based large-scale parallel relation-learning system
  • Extracted feeds form social media sites such as Facebook, Twitter using Flume,Python scripts.
  • Setup and benchmarked Hadoop/HBase clusters for internal use

Environment: Hadoop, MapReduce, HDFS, Hive, Java, Pig, HBase, Linux, XML, Java 6, Eclipse, Oracle 10g, PL/SQL, MongoDB, Toad

Confidential, Houston, TX

Oracle PL/SQL Developer

Responsibilities:

  • Worked as a lead developer for TCOA Volume increase and COMIT Projects.
  • Gathering requirements from Business partners and writing technical designs and hand over to the development team.
  • Managing Cost Center collapses every month and running migration scripts.
  • While merging data from TCOA, TCO applications, impact analysis for the volume increase.
  • Analyze code and document Technical designs, Data Model diagrams, Process flow diagrams, Data flow diagrams as per industry documentation standards.
  • Created Table partitions as per the DATASET, to improve the performance.
  • Working on all the environments Like UAT, SIT, PROD environments.
  • Working on Performance tuning and Efficient Coding techniques.
  • Created dynamic parameter files scripts to extract Data.
  • Responsible for monitoring scheduled, running completed and failed sessions. Involved in debugging the failed mappings and developing error-handling methods.
  • Involved in writing stored procedures and Shell Scripts for automating the execution of sessions.
  • Attending in the Functional Design Meetings and gathering requirements of the application.
  • Extensively Participated in Impact Analysis for the new requirements.
  • Create Technical Designs for the Functional designs and making DDL Script if necessary.
  • Developing Test files for the application such that every possible scenario works well efficiently.
  • Optimized the queries to improve the performance of the application using Explain plan and TkProf.
  • Modifying the existing technical specifications to be adjusted to the new user requirements.
  • Writing PL/SQL code using the technical and functional specifications.
  • Document all the changes and Check in the code in Source safe in time with the defect numbers.
  • DBMS utilities were used to extend the functionality of PL/SQL programs such as Execute immediate for scheduling, UTL FILE to read and write from database to generate data file, to write Dynamic SQL etc.
  • Designed Oracle Objects including Triggers, Stored Procedures, and Sequences that helped in automating various operations in the database.
  • Created indexes on tables to improve performance.
  • Extensively used Explicit Cursors in stored procedures, Dynamic SQL, Analytic Functions and global temporary tables for intermediate processing of data.

Environment: Oracle 11g, SQL * Plus, SQL Navigator, TOAD, Tortoise Subversion, Microsoft Visio, Microsoft Project, Test Director, Rational Clear case, Rational Clear Quest, Unix, Shell Scripting, Microsoft Office.

Confidential

Oracle PL/SQL Developer

Responsibilities:

  • Designed and developed user interfaces using JSP, Java script and HTML.
  • Developed web components using JSP, Servlets and JDBC
  • Used JDBC and managed connectivity, for inserting/querying& data management including stored procedures and triggers.
  • Developed Database applications using SQL and PL/SQL.
  • Involved in the design and coding of the data capture templates, presentation and component templates.
  • Developed an API to write XML documents from database.
  • Involved in fixing bugs and unit testing with test cases using Junit
  • Part of a team which is responsible for metadata maintenance and synchronization of data from database.

Environment: Java script, JSP, JDBC, Servlets, HTML, XML, SQL, Junit.

Confidential

Oracle PL/SQL Developer

Responsibilities:

  • Generated DDL statements for creation of new database objects like tables, views, sequences, functions, synonyms, indexes, triggers, packages and stored procedures.
  • Developed Database Triggers to enforce security also used ref cursors.
  • Generated server side PL/SQL scripts for data manipulation and validation and created various materialized views for remote instances.
  • Modified existing stored procedures in order to implement the suggested changes in the system
  • Involved in Performance Optimization to increase operational efficiency Using SQL Trace,
  • TKPROF and Explain Plan Utilities.
  • Developed data mapping spreadsheet to present the transformation process.
  • Involved in writing stored procedures
  • Expertise in performing optimization analysis.
  • Optimized the Cache for Dynamic, Static and Persistent Cache Lookup Transformations.
  • Tuned performance of Informatics session for large data files by increasing data cache size and target based commit interval. Created New Messages and Alerts as part of Forms Customization.
  • Used Perlto write command line scripts and to parse text.
  • Developed User Documentation for all Application Modules following stringent standards.

Environment: Oracle 11g, SQL Developer, Oracle forms and reports developer, SQL*Loader, SQL plus.

Confidential

Jr. Oracle Developer

Responsibilities:

  • Analyzed the business requirements of the project by studying the Business Requirement Specification Document.
  • Creation of database objects like tables, views, procedures, packages using Oracle utilities like SQL Plus, SQL*Loader and Exception handling.
  • Timely monitoring of database and system backups developed Oracle stored procedures, functions, packages and triggers that pull data for reports.
  • Designed the front end interface for the users using Oracle Forms, Reports using 8i.
  • Involved in database development by creating Oracle PL/SQL functions, Procedure, Triggers, Packages, Records and Collections.
  • Run batch jobs for loading database tables from Flat files and CSV using SQL Loader, Data mapping.
  • Participated in Performance Tuning for query optimization and program performance.
  • Created UNIX shell scripts for automating the execution process.
  • Automated quality check tasks by creating PL/SQL Cursors, Functions and Dynamic SQL.

Environment: Oracle 9i, UNIX, ERWIN, SQL Developer, Oracle forms and reports developer, SQL*Loader, SQL plus.

We'd love your feedback!