We provide IT Staff Augmentation Services!

Hadoop Developer Resume

0/5 (Submit Your Rating)

Chicago, IL

SUMMARY:

  • Over 7+ years of experience in software development, deployment and maintenance of web - based applications using Big Data Ecosystems on various environments.
  • Experienced in using Agile methodologies including extreme programming, SCRUM and Test-Driven Development (TDD).
  • Expertise in major components of Hadoop ecosystems like HDFS, MapReduce, YARN, Hive, Pig, HBase, Zookeeper, Sqoop, Spark, Kafka, Cassandra and Impala.
  • Hands-on experience on installing, configuring and maintaining multi-node clusters on various environments and distributions of Hadoop.
  • Experience in importing and exporting different formats of data into HDFS, HBASE from different RDBMS databases and vice versa using Sqoop.
  • Expertise in implementing Spark and Scala application using higher order functions for both batch and interactive analysis requirement.
  • Excellent experience in AWS, Cloudera maintaining and optimized AWS infrastructure (EC2 and EBS) also good knowledge in MS Azure.
  • Hands on experience in configuring and working with Flume to load the data from multiple sources directly into HDFS.
  • Experience in developing custom UDF's for Pig and Apache Hive to in corporate methods and functionality of Java into Pig Latin and HiveQL.
  • Experienced in job workflow scheduling & monitoring tools like Oozie and Zookeeper.
  • Expertise in developing production ready Spark applications utilizing Spark-Core, Data frames, Spark-SQL, Spark-ML and Spark-Streaming API's.
  • Exposure to administrative tasks such as installing Hadoop and its ecosystem components such as Hive and Pig.
  • Expertise in using ETL Tool Informatica PowerCenter designer, workflow manager, repository manager, data quality and ETL concepts.
  • Experience in creating complex SQL Queries and SQL tuning, writing PL/SQL blocks like stored procedures, Functions, Cursors, Index, triggers and packages.
  • Committed to professionalism, highly organized, ability to work under strict deadline schedules with attention to details, possess excellent written and communication skills.

TECHNICAL SKILLS

Hadoop: HDFS, MapReduce, YARN, Hive, Pig, HBASE, Impala, Zookeeper, Sqoop, OOZIE, Apache Cassandra, Flume, Spark, AWS, EC2

Languages: C, Java, SQL, PL/SQL, Scala, Shell Scripts

Operating Systems: Linux, UNIX, Windows

Databases: Oracle, SQL Server, Teradata, MS Access, HBase

Application Servers: WebLogic, WebSphere, Apache Tomcat, JBOSS

IDE’s: Eclipse, NetBeans JDeveloper, IntelliJ IDEA.

Version Control: CVS, SVN, GIT

Web Technologies: HTML, CSS, JavaScript, AJAX, Servlets, JSP, DOM, XML, XSLT.

PROFESSIONAL EXPERIENCE:

Confidential, Chicago, IL

Hadoop Developer

Responsibilities:

  • Coordinated with business customers to gather business requirements and interacted with other technical peers to derive technical requirements.
  • Involved in story-driven Agile development methodology and actively participated in daily Scrum meetings.
  • Wrote Map Reduce code that will take input as customer related flat file and parse the same data to extract the meaningful (domain specific) information for further processing.
  • Worked on creating Combiners, Partitioning and Distributed cache to improve the performance of Map Reduce jobs.
  • Performed performance tuning and troubleshooting of MapReduce jobs by analyzing and reviewing Hadoop log files.
  • Created Hive tables to import large data sets from relational databases using Sqoop and export analyzed data back for visualization and report generation by BI team.
  • Developed Hive Scripts to create views & apply transformation logic in Target Database.
  • Involved in developing Hadoop Map Reduce jobs using Java Environment.
  • Involved in design of Data Mart and Data Lake to provide faster insight into the Data.
  • Developed Pig Latin scripts to extract data from server output files to load in HDFS.
  • Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS.
  • Involved in using Stream Sets Data Collector tool and created Data Flows for one of the streaming applications.
  • Worked on AWS provisioning EC2 Infrastructure and deploying applications in Elastic load balancing.
  • Implemented Kafka High level consumers to get data from Kafka partitions and move into HDFS.
  • Developed a script in Scala to read all the Parquet Tables in a Database and parse them as JSON files, another script to parse them as structured tables in Hive.
  • Performed various operations on data lake of MapR cluster which involves moving the data, enriching the data, performing validations.
  • Data analysis on use cases, Customer communication, Incident management, Production support.
  • Maintained MapReduce jobs to ensure maintenance of the MapR Data lake Cluster.
  • Imported data from different sources like AWS S3, Local file system into Spark RDD.
  • Developed and implemented a Unix Shell script which retrieves the metadata of all the hive tables in a database.

Confidential, Birmingham, AL

Hadoop Developer

Responsibilities:

  • Involved in installation, configuration, design, development, and maintenance of complete SDLC in an Agile methodology.
  • Prepared Spark builds from MapReduce source code for better performance.
  • Involved in importing the real-time data to Hadoop using Kafka and implemented Oozie jobs for daily imports.
  • Worked on Oozie workflow engine for job scheduling Imported and exported data into MapReduce and Hive using Sqoop.
  • Configured AWS RDS/Redshift to use Hadoop Ecosystem on AWS infrastructure.
  • Developed Hive scripts for performing transformation logic and also loading the data from staging zone to final landing zone.
  • Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
  • Developed Scripts and Batch Job to schedule various Hadoop Program.
  • Implemented NiFi flow topologies to perform cleansing operations before moving data into HDFS.
  • Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS.
  • Used Spark API over Hortonworks, Hadoop YARN to perform data analysis in Hive.
  • Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi structured data coming from various sources.
  • Developed a shell script to create staging, landing tables with the same schema like the source and generate the properties which are used by Oozie jobs.
  • Working on cluster co-ordination with data capacity planning and node forecasting using Zookeepers.
  • Worked on developing ETL Workflows on the data obtained using Python for processing it in HDFS and HBase using Oozie.
  • Involved on configuration, development of Hadoop environment with AWS cloud such as EC2, EMR, Redshift, Route 53, Cloud watch.
  • Use Flume to aggregate and store the web log data obtained from various sources such as web servers and network devices.
  • Involved in developing Scala programs which supports functional programming.
  • Used GIT-Hub for project version management.

Confidential, Omaha, NE

Hadoop Developer

Responsibilities:

  • Responsible for generating actionable insights from complex data to drive significant business results for various application teams.
  • Troubleshoot and resolve data quality issues and maintain important level of data accuracy in the data being reported.
  • Worked with application teams to install Hadoop updates, patches & version upgrades.
  • Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
  • Implemented best income logic using Pig scripts and UDFs.
  • Implemented test scripts to support test driven development and continuous integration.
  • Managed and reviewed Hadoop log files. Used Scala for integration Spark into Hadoop.
  • Responsible to manage data coming from different sources. Implemented Oozie workflow engine to run multiple Hive and Python jobs.
  • Migrated data existing in Hadoop cluster into Spark and used SparkSQL and Scala to perform actions on the data.
  • Created Hive Partitions for storing data for different trends under different partitions.
  • Connected the Hive tables to data analysis tools like Tableau for graphical representation of the trends.
  • Assisted project manager in problem shooting relevant to Hadoop technologies for data integration between different platforms like Sqoop-Sqoop, Hive-Sqoop, and Sqoop-Hive.
  • Designed ETL process using Teradata to load the data from various source databases and flat files to target data warehouse in Oracle
  • Used Power mart Workflow Manager to design sessions, event wait/raise, and assignment, e-mail, and command to execute mappings
  • Created parameter-based mappings, Router and lookup transformations
  • Involved in migration projects to migrate data from data warehouses on Oracle and migrated those to Teradata
  • Optimized mappings using transformation features like Aggregator, Filter, Joiner, Expression and Lookups
  • Implemented J2EE design patterns such as singleton, DAO for the presentation tier, business tier and Integration Tier layers of the project.
  • Involved in Bug fixing and functionality enhancements.
  • Followed coding and documentation standards and best practices.

Confidential, Jackson, MI

Java Developer

Responsibilities:

  • Responsible and active in the analysis, design, implementation and deployment of SDLC of the project.
  • Designed Use Case Diagrams, Class Diagrams and Sequence Diagrams and Object Diagrams to model the detail design of the application using UML.
  • Developed the DAO layer using the Hibernate annotations and configuration files.
  • Used Spring MVC Framework Dependency Injection for integrating various Java Components.
  • Designed and developed user interface using JSP, HTML and JavaScript.
  • Validated the fields of user registration screen and login screen by writing JavaScript and jQuery validations.
  • Configured the Spring framework for the entire business logic layer.
  • Wrote the Map Reduce jobs to parse the web logs which are stored in HDFS.
  • Developed Simple to complex MapReduce Jobs using Hive and Pig.
  • Involved in testing and deployment of the application on WebSphere Application Server during integration and QA testing phase.
  • Used Maven Scripts to build and deploy applications and helped to deployment for Continuous Integration using Jenkin and Maven.
  • Wrote SQL queries and Stored Procedures for interacting with the Oracle database.
  • Use the XML based request and response messages for communication and uses the DTDs for validation.
  • Developed the Message Driven Beans for purging utilities of audit log tables using JMS services.
  • Worked on the presentation and UI components using XSL, CSS and JavaScript with Builder design pattern.
  • Used Log4J for logging framework to debug the code.
  • Documentation of common problems prior to go-live and while actively in a Production Support role.

We'd love your feedback!