Sr. Hadoop Developer Resume
Reston, VA
SUMMARY
- Overall 8+ years of IT experience in the field of Information Technology that includes analysis, design, development and testing of complex applications.
- Around 2+ years of strong working experience with Big Data and Hadoop Ecosystems.
- Strong experience with configuring, using Hadoop and Ecosystem components like HBase, Hive, Pig, Oozie, Sqoop, Flume, Mahout, Impala.
- Excellent understanding / knowledge of Hadoop architecture and various components of Hadoop such as HDFS, JobTracker, TaskTracker, NameNode, DataNode, Map Reduce & YARN.
- Working Experience in writing Map Reduce Programs using Java and R.
- Worked with HiveQL to query data from Hive tables in HDFS.
- Used Pig Latin scripts and customized UDF’s in Java to analyze large data sets.
- Experience in using Hadoop distributions like ClouderaCDH4, Hortonworks2.0, Amazon Web Services and open source Apache Hadoop.
- Extensive exposure in RDBMS architecture, modelling, design, development, loading migration with Oracle, SQL Server, DB2 and experience with NOSQL data stores like HBase, MongoDB.
- Extensive experience in developing PL/SQL programming structures such as Stored Procedures, Functions, Packages, Views, Cursors and Triggers.
- Experience in data ingress and egress using Sqoop from HDFS to Relational Database Systems and vice - versa.Good knowledge of Log4j for error handling.
- Experience in working with Flume to load the log data from multiple sources directly into HDFS.
- Experience in designing both time driven and data driven automated workflows using Oozie and configuring ZooKeeper for distributed synchronization and providing group services.
- WorkingExperience in writing Pythonand RMapReduce Scripts to leverage Hadoop Streaming.
- Experience in processing data serialization formats like Xml, JSON, SequenceFiles and Avro.
- Developed Unit tests for MapReduce code using MRUnit test framework.
- Good understanding in LZOP and Snappy Compression Codec techniques with PARQUET column format
- Developed few comprehensive requirements & POC’s to understand Mahout in Data science for Predictive analytics.
- Good experience in understanding and visualizing data by integrating Hadoopwith Tableau and analytics using Splunk.
- Good experience in diagnosing, tuning, profiling map and reduce tasks.
- Good understanding of star schema, snowflake schema, Change Data Capture, Kimball and Inmon methodologies.
- Worked on Informatica Power Center 9.1/8.6.1/8.5.1/8.1/7.1/6.1 , Talend for ETL and Report generation using Business Objects.
- Extensive experience in middle-tier development using J2EE technologies like JDBC, JNDI, JSP, Servlets, JSP, JSF, Spring, Hibernate, Struts, JDBC, EJB.
- Experience with web-based UI development using jQuery, CSS, HTML, HTML5, XHTML and JavaScript.
- Proficient in all the phases of Software Development Cycle (SDLC) and implemented various methodologies like Waterfall, Spiral, Agile SCRUM.
- Ability to learn and master new technologies, deliver outputs in short deadlines and having good programming skills.
- Excellent with coordinating, mentoring and exchanging information with colleagues.
TECHNICAL SKILLS
Big Data& Ecosystem: HDFS, Map Reduce, HBase, Pig, Hive, Sqoop, Flume, Mahout, MongoDB, Oozie, Tez, Impala, SOLR
Big data Analytics: Splunk, Tableau
Java & J2EE Technologies: Core Java, Servlets, JSP, JDBC, JNDI, Java Beans, Log4j
Frameworks: MVC, Struts, Hibernate, Spring
Programming languages: C, C++, Java, Python, R, Linux Shell Script, SQL, HiveQL, Pig Latin
IDE’s: Eclipse, Net beans, RStudio, Visual Studio 2012
Databases: Oracle 11g/10g/9i, MySQL, DB2, Teradata
Application Servers: Apache Tomcat 5/6/7, JBoss 6/7
Web Technologies: HTML, XML, JavaScript, SOAP, DOM
ETL Tools: Talend, Informatica PowerCenter 9.x/8.x
Testing: MRUnit, Load Runner, JUnit
Operating Systems: Windows, Ubuntu 12.04, CentOS 5.x, RHEL
Virtualization Environment: Oracle VB, VMware, Amazon EC2
Source Code Management: Git
Others: Weka 3, Revolution R, winSCP, Putty, MS-Office, Oracle BI
PROFESSIONAL EXPERIENCE
Confidential, Reston, VA
Sr. Hadoop Developer
Responsibilities:
- As a Functional Lead Involved in design, development and implementation of Predictive analytic models.
- Wrote Map Reduce jobs to launch and monitor processing-intensive computations on Cloudera.
- Responsible for data ingress and egress using Sqoop from HDFS to databases and vice-versa
- Wrote Pig Latin Scripts to generate MapReduce jobs and performed transformation procedures on the data in HDFS.
- Processed HDFS data and created external tables using HiveQL by partitioning and bucketing for granularity in order to analyze spikes, faults, optimize for customer experience.
- Developed enhanced model using the clustering techniques with Mahout Machine learning library
- Developed few requirements and POC’s to achieve interactive query for Hadoop with Apache Tez.
- Involved in identifying job dependencies to design workflow for Oozie and resource management for YARN.
- Worked with common serialization formats like Json, Xml and Big data serialization formats like Avro and SequenceFiles.
- AppliedHunkanalytics platform to analyze, visualize data and create reports for analysis.
- Providing support for deployment of code to multiple environments.
- Used Python Scipy packages and scripts to leverage Hadoop Streaming.
- Developed few requirements and Proof of concepts to integrate R on top of the Hadoop platform (Rhadoop, R+ streaming) to perform data analysis and graphing the results.
- R has applicability across wide range of sectors and makes it a powerful tool.
Environment: Hadoop, MapReduce, HDFS, Amazon EC2, Amazon EBS, Hive, HBase, Cloudera, Java (Jdk 1.6), Eclipse, Linux, Tez, Mahout, Pig, Storm, Python, R, Cassandra, Talend
Confidential
Sr. Hadoop Developer
Responsibilities:
- Analyzed business requirements and existing software for High Level Design.
- Designed detailed software structure and architecture documents using Use cases, sequence diagram and UML.
- Experience in working with Flume to load the log data from multiple sources directly into HDFS.
- Developed Unit tests for Map Reduce code using MRUnit test framework.
- Applied Map Reduce patterns like repartition-joins, semi-joins and sampling to big data.
- Used Stack dumps to discover unoptimized user code.
- Used Crunch log parsing to find the URL patterns and basic analytics.
- Written Sqoop incremental pooling job to move new / updated info from Database to HDFS and HBase.
- Applied Tableau analytics platform to analyze, visualize data and create reports for analysis.
- Created Oozie coordinated workflow to execute Sqoop incremental job daily.
- Examined Task logs and figured out the JVM startup arguments for a task
- Developed POC’s for Splittable LZOP, Snappy and LZOP Compression Codec techniques.
Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Java, SQL Server, Cloudera Manager, Sqoop, Flume, Oozie, Java (Jdk 1.6), Eclipse, Linux, Tableau, MongoDB, Talend
Confidential
ETL Developer
Responsibilities:
- Collaborated with Business analysts and the DBA for requirements gathering, business analysis and designing of the data marts.
- Designed, developed Informaticav8.1.1mappings, to extract, transform and loading of the data into Oracle 10g target tables.
- Designed and Developed Source to Stage Mappings/Sessions/Workflows.
- Designed and Developed Stage to ODS Mappings/Sessions/Workflows.
- Designed and Developed ODS to DataMart Mappings/Sessions/Workflows.
- Worked with heterogeneous sources from various channels like Oracle, SQL Server, flat files.
- Created User, Groups, assigning users to groups and grant privileges and permissions to groups and folders.
- Worked on Informatica tool Source Analyzer, Warehouse Designer, Mapping Designer, Workflow Manager, Workflow Monitor, and Repository Manager.
- Extensively used Transformations like, Aggregator, Router, Joiner, Expression, Lookup, Update Strategy, and Sequence Generator.
- Used Informatica Repository Manager to create Repositories and Users and to give permissions to users.
- Developed UNIX Shell Scripts and PL/SQL procedures.
- Extensively worked in the performance tuning of ETL mappings and sessions.
- Used TOAD to run SQL queries and validate the data in warehouse and mart.
- Tested the data and data integrity among various sources and targets. Associated with Production support team in various performances related issues.
Environment: Informatica Power Center 8.1.1, Oracle 10g, DB2, PL/SQL, SQL Server 2005, Toad, Windows XP, Unix Shell Scripts, Erwin.
Confidential
Jr Software Developer
Responsibilities:
- Involved in the Analysis of the requirement specifications provided by the client.
- Development, Testing and Code review of the application.
- Used JDBC to connect the J2EE server with the relational database.
- User input validations done using JavaScript
- Extreme programming methodologies for replacing the existing code and testing in J2EE environment.
- Developed use cases using UML and java classes for business layer.
- Transactions boundaries are defined with annotations provided by EJB specification.
- Developed XML schemas for defining request and response messages for web services.
- Used Soap UI tool for testing and monitoring soap request and response.
- Designed and developed web based application using JSF framework for implementing MVC and component based architecture.
- Worked in developing EJB3’s session bean and entity beans for implementing business logics.
- Maintained stored definitions, transformation rules and targets definitions using Informatica repository Manager.
- Used various transformations like Filter, Expression, Sequence Generator, Update Strategy, Joiner, Stored Procedure, and Union to develop robust mappings in the Informatica Designer.
- Used Type 1 SCD and Type 2 SCD mappings to update slowly Changing Dimension Tables.
- Developed database objects like tables, views, stored procedures, indexes.
- Involved in the enhancement of some applications and user requirements (Change Requests).
- Facilitate communication within the project team.
Environment: Core Java, HTML, JavaScript, JDBC, MYSQL, JSF, PL/SQL, EJB, Microsoft Visio, Eclipse 3.2, Informatica Power Center 8.6.1, Workflow Manager, PL/SQL, Oracle 10g/9i, Erwin.
