Hadoop Developer Resume
Chicago, IL
SUMMARY:
- Over 7+ years of experience in software development, deployment and maintenance of web - based applications using Big Data Ecosystems on various environments.
- Experienced in using Agile methodologies including extreme programming, SCRUM and Test-Driven Development (TDD).
- Expertise in major components of Hadoop ecosystems like HDFS, MapReduce, YARN, Hive, Pig, HBase, Zookeeper, Sqoop, Spark, Kafka, Cassandra and Impala.
- Hands-on experience on installing, configuring and maintaining multi-node clusters on various environments and distributions of Hadoop.
- Experience in importing and exporting different formats of data into HDFS, HBASE from different RDBMS databases and vice versa using Sqoop.
- Expertise in implementing Spark and Scala application using higher order functions for both batch and interactive analysis requirement.
- Excellent experience in AWS, Cloudera maintaining and optimized AWS infrastructure (EC2 and EBS) also good knowledge in MS Azure.
- Hands on experience in configuring and working with Flume to load the data from multiple sources directly into HDFS.
- Experience in developing custom UDF's for Pig and Apache Hive to in corporate methods and functionality of Java into Pig Latin and HiveQL.
- Experienced in job workflow scheduling & monitoring tools like Oozie and Zookeeper.
- Expertise in developing production ready Spark applications utilizing Spark-Core, Data frames, Spark-SQL, Spark-ML and Spark-Streaming API's.
- Exposure to administrative tasks such as installing Hadoop and its ecosystem components such as Hive and Pig.
- Expertise in using ETL Tool Informatica PowerCenter designer, workflow manager, repository manager, data quality and ETL concepts.
- Experience in creating complex SQL Queries and SQL tuning, writing PL/SQL blocks like stored procedures, Functions, Cursors, Index, triggers and packages.
- Committed to professionalism, highly organized, ability to work under strict deadline schedules with attention to details, possess excellent written and communication skills.
TECHNICAL SKILLS
Hadoop: HDFS, MapReduce, YARN, Hive, Pig, HBASE, Impala, Zookeeper, Sqoop, OOZIE, Apache Cassandra, Flume, Spark, AWS, EC2
Languages: C, Java, SQL, PL/SQL, Scala, Shell Scripts
Operating Systems: Linux, UNIX, Windows
Databases: Oracle, SQL Server, Teradata, MS Access, HBase
Application Servers: WebLogic, WebSphere, Apache Tomcat, JBOSS
IDE’s: Eclipse, NetBeans JDeveloper, IntelliJ IDEA.
Version Control: CVS, SVN, GIT
Web Technologies: HTML, CSS, JavaScript, AJAX, Servlets, JSP, DOM, XML, XSLT.
PROFESSIONAL EXPERIENCE:
Confidential, Chicago, IL
Hadoop Developer
Responsibilities:
- Coordinated with business customers to gather business requirements and interacted with other technical peers to derive technical requirements.
- Involved in story-driven Agile development methodology and actively participated in daily Scrum meetings.
- Wrote Map Reduce code that will take input as customer related flat file and parse the same data to extract the meaningful (domain specific) information for further processing.
- Worked on creating Combiners, Partitioning and Distributed cache to improve the performance of Map Reduce jobs.
- Performed performance tuning and troubleshooting of MapReduce jobs by analyzing and reviewing Hadoop log files.
- Created Hive tables to import large data sets from relational databases using Sqoop and export analyzed data back for visualization and report generation by BI team.
- Developed Hive Scripts to create views & apply transformation logic in Target Database.
- Involved in developing Hadoop Map Reduce jobs using Java Environment.
- Involved in design of Data Mart and Data Lake to provide faster insight into the Data.
- Developed Pig Latin scripts to extract data from server output files to load in HDFS.
- Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS.
- Involved in using Stream Sets Data Collector tool and created Data Flows for one of the streaming applications.
- Worked on AWS provisioning EC2 Infrastructure and deploying applications in Elastic load balancing.
- Implemented Kafka High level consumers to get data from Kafka partitions and move into HDFS.
- Developed a script in Scala to read all the Parquet Tables in a Database and parse them as JSON files, another script to parse them as structured tables in Hive.
- Performed various operations on data lake of MapR cluster which involves moving the data, enriching the data, performing validations.
- Data analysis on use cases, Customer communication, Incident management, Production support.
- Maintained MapReduce jobs to ensure maintenance of the MapR Data lake Cluster.
- Imported data from different sources like AWS S3, Local file system into Spark RDD.
- Developed and implemented a Unix Shell script which retrieves the metadata of all the hive tables in a database.
Confidential, Birmingham, AL
Hadoop Developer
Responsibilities:
- Involved in installation, configuration, design, development, and maintenance of complete SDLC in an Agile methodology.
- Prepared Spark builds from MapReduce source code for better performance.
- Involved in importing the real-time data to Hadoop using Kafka and implemented Oozie jobs for daily imports.
- Worked on Oozie workflow engine for job scheduling Imported and exported data into MapReduce and Hive using Sqoop.
- Configured AWS RDS/Redshift to use Hadoop Ecosystem on AWS infrastructure.
- Developed Hive scripts for performing transformation logic and also loading the data from staging zone to final landing zone.
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis.
- Developed Scripts and Batch Job to schedule various Hadoop Program.
- Implemented NiFi flow topologies to perform cleansing operations before moving data into HDFS.
- Involved in file movements between HDFS and AWS S3 and extensively worked with S3 bucket in AWS.
- Used Spark API over Hortonworks, Hadoop YARN to perform data analysis in Hive.
- Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi structured data coming from various sources.
- Developed a shell script to create staging, landing tables with the same schema like the source and generate the properties which are used by Oozie jobs.
- Working on cluster co-ordination with data capacity planning and node forecasting using Zookeepers.
- Worked on developing ETL Workflows on the data obtained using Python for processing it in HDFS and HBase using Oozie.
- Involved on configuration, development of Hadoop environment with AWS cloud such as EC2, EMR, Redshift, Route 53, Cloud watch.
- Use Flume to aggregate and store the web log data obtained from various sources such as web servers and network devices.
- Involved in developing Scala programs which supports functional programming.
- Used GIT-Hub for project version management.
Confidential, Omaha, NE
Hadoop Developer
Responsibilities:
- Responsible for generating actionable insights from complex data to drive significant business results for various application teams.
- Troubleshoot and resolve data quality issues and maintain important level of data accuracy in the data being reported.
- Worked with application teams to install Hadoop updates, patches & version upgrades.
- Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
- Implemented best income logic using Pig scripts and UDFs.
- Implemented test scripts to support test driven development and continuous integration.
- Managed and reviewed Hadoop log files. Used Scala for integration Spark into Hadoop.
- Responsible to manage data coming from different sources. Implemented Oozie workflow engine to run multiple Hive and Python jobs.
- Migrated data existing in Hadoop cluster into Spark and used SparkSQL and Scala to perform actions on the data.
- Created Hive Partitions for storing data for different trends under different partitions.
- Connected the Hive tables to data analysis tools like Tableau for graphical representation of the trends.
- Assisted project manager in problem shooting relevant to Hadoop technologies for data integration between different platforms like Sqoop-Sqoop, Hive-Sqoop, and Sqoop-Hive.
- Designed ETL process using Teradata to load the data from various source databases and flat files to target data warehouse in Oracle
- Used Power mart Workflow Manager to design sessions, event wait/raise, and assignment, e-mail, and command to execute mappings
- Created parameter-based mappings, Router and lookup transformations
- Involved in migration projects to migrate data from data warehouses on Oracle and migrated those to Teradata
- Optimized mappings using transformation features like Aggregator, Filter, Joiner, Expression and Lookups
- Implemented J2EE design patterns such as singleton, DAO for the presentation tier, business tier and Integration Tier layers of the project.
- Involved in Bug fixing and functionality enhancements.
- Followed coding and documentation standards and best practices.
Confidential, Jackson, MI
Java Developer
Responsibilities:
- Responsible and active in the analysis, design, implementation and deployment of SDLC of the project.
- Designed Use Case Diagrams, Class Diagrams and Sequence Diagrams and Object Diagrams to model the detail design of the application using UML.
- Developed the DAO layer using the Hibernate annotations and configuration files.
- Used Spring MVC Framework Dependency Injection for integrating various Java Components.
- Designed and developed user interface using JSP, HTML and JavaScript.
- Validated the fields of user registration screen and login screen by writing JavaScript and jQuery validations.
- Configured the Spring framework for the entire business logic layer.
- Wrote the Map Reduce jobs to parse the web logs which are stored in HDFS.
- Developed Simple to complex MapReduce Jobs using Hive and Pig.
- Involved in testing and deployment of the application on WebSphere Application Server during integration and QA testing phase.
- Used Maven Scripts to build and deploy applications and helped to deployment for Continuous Integration using Jenkin and Maven.
- Wrote SQL queries and Stored Procedures for interacting with the Oracle database.
- Use the XML based request and response messages for communication and uses the DTDs for validation.
- Developed the Message Driven Beans for purging utilities of audit log tables using JMS services.
- Worked on the presentation and UI components using XSL, CSS and JavaScript with Builder design pattern.
- Used Log4J for logging framework to debug the code.
- Documentation of common problems prior to go-live and while actively in a Production Support role.
