Etl - Hadoop Architect/developer Resume
NJ
SUMMARY
- Over 8 years of extensive software development experience in full life cycle development which includes more than 3 years of experience as Hadoop Developer/Architect focusing on various Big Data Technologies.
- Experience in developing Map Reduce Programs using Apache Hadoop for analyzing teh big data as per teh requirement.
- Hands on experience in writing MapReduce jobs in Java, Pig and Python.
- Developing both batch and real time applications on teh Hadoop platform.
- Experienced on major Hadoop ecosystem’s projects such as PIG, HIVE and HBASE.
- Good working experience using Sqoop to import data into HDFS from RDBMS and vice - versa.
- Good knowledge in using job scheduling and monitoring tools like Oozie and ZooKeeper.
- Knowledge of NoSQL databases such as HBase, and MongoDB.
- Good understanding/knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
- Has a good understanding to ETL concepts.
- Hands on experience in installing, configuring and using ecosystem components like HadoopMapReduce, HDFS, Hbase, ZooKeeper, Oozie, Flume, Sqoop, Pig & Hive with CDH3&4&5 clusters..
- Experience in managing Hadoop clusters using Cloudera Manager
- Experience in database development using SQL and PL/SQL and experience working on databases like Oracle 9i/10g, Informix, and SQL Server.
- Performeddata analysisusingMySQL, SQL Server Management Studio, and Oracle.
- Profound knowledge of teh principal of DW using Fact Tables, Dimension Tables, Star schema modeling and Snowflake Schema modeling.
- Strong skills in Datastage Administrator in UNIX & LINUX environments, Report creation using OLAP data source and having knowledge in OLAP universe.
- Worked on debugging tools such as Dtrace, Struss and Top. Expert in setting up SSH, SCP, SFTP connectivity between UNIX hosts.
- Experienced teh integration of various data sources like Java, RDBMS, Shell Scripting, Spreadsheets, and Text files.
- Used Springs JDBC and DAO layers to offer abstraction for teh business from teh database related code (CRUD).
- Working with relative ease with different working strategies like Agile, Waterfall and Scrum methodologies.
- Expertise in designing and developing J2EE compliant systems using IDE tools like Eclipse, WebSphere Studio Application Developer (WSAD).
- In-depth understanding of Data Structures and Algorithms.
- UNIX shell scripting, resource Extensive experience in Java and J2EE technologies like Servlets, JSP, JSF, and JDBC.
- Expert in using J2EE complaint application servers Apache Tomcat, IBM Web Sphere.
- Extensively worked ondebuggingusing Eclipse debugger.
- Hands on experience working with Java project build managers Apache MAVEN and ANT
- Good Knowledge in Flume, Avro and Zoo Keeper Architecture.
- Good understanding of Data Mining and Machine Learning techniques.
- Extensive experience in working with different databases such as Oracle, IBM DB, RDBMS, SQL Server, MySQL and writing Stored Procedures, Functions, Joins and Triggers for different Data Models.
- Extensive experience in working with Parallel jobs, Troubleshooting and Performance tuning.
- Strong work ethic with desire to succeed and make significant contributions to teh organization.
- Strong problem solving skills, good communication, interpersonal skills and a good team player.
- Has teh motivation to take independent responsibility as well as ability to contribute and be a productive team member.
TECHNICAL SKILLS
Hadoop/Big Data Technologies: HDFS, Map Reduce, Hive, Pig, Sqoop, Flume, Oozie, and Zookeeper.
No SQL Databases: Hbase,Cassandra, mongoDB.
Languages: C, C++, Java, J2EE, C#, Asp.Net, PL/SQL, Pig Latin, HiveQL, Unix shell Scripts.
Java/J2EE Technologies: Applets, Swing, JDBC, JNDI, JSON, JSTL, RMI, JMS, Java Script, JSPServlets, EJB, JSF, JQuery.
Frameworks: MVC, Struts, Spring, And Hibernate.
Web Services: HTML, DHTML, XML, AJAX, WSDL, SOAP, REST, Jersey.
Web/Application servers: Apache Tomcat, WebLogic, JBoss, Web Sphere.
Operating Systems: Sun Solaris, HP-UNIX, RedHat Linux, Ubuntu Linux and Windows XP/Vista/7/8.
Databases: Oracle 9i/10g/11g, DB2, SQL Server, MySQL.
Web technologies: JSP, Servlets, JNDI, JDBC, Java Beans, JavaScript, Web Services (JAX-WS).
Java IDE: Eclipse 3.x, IBM Web Sphere Application Developer, IBM RAD 7.0.
Tools: TOAD, SQL Developer, SOAP UI, ANT, Maven, Visio, Rational RoseEndur 8.x/10.x/11.x.
PROFESSIONAL EXPERIENCE
Confidential, NJ
ETL - Hadoop Architect/Developer
Responsibilities:
- Responsible for building scalable distributed data solutions using Hadoop.
- Installed and configured Hive, Pig, Sqoop, Flume and Oozie on teh Hadoop cluster.
- Developed Simple to complex Map/reduce Jobs using Hive and Pig.
- Devised and lead teh implementation of teh next generation architecture for more efficient data ingestion and processing.
- Give extensive presentations about teh Hadoop ecosystem, best practices, data architecture in Hadoop.
- Provide mentorship and guidance to other architects to halp them become independent.
- Provide review and feedback for existing physical architecture, data architecture and individual code.
- Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms.
- Involved in loading data from UNIX file system to HDFS.
- Wrote MapReduce jobs to discover trends in data usage by users.
- Involved in running Hadoop streaming jobs to process terabytes of text data.
- Analyzed large data sets by running Hive queries and Pig scripts.
- Helped teh team to increase teh Cluster size from 22 to 30 Nodes.
- Job management using Fair scheduler.
- Develop Core Framework based on Hadoop to Migrate Existing ETL (RDBMS) Solution.
- Wrote Pig Scripts to generate Map Reduce jobs and performed ETL procedures on teh data in HDFS
- A deep and thorough understanding of ETL tools and how they can be applied in a Big Data environment.
- Worked as ETL Architect to make sure all teh applications are migrated (along with server) smoothly.
- Involved in creating Hive tables, and loading and analyzing data using hive queries.
- Responsible for managing data from multiple source.
- Designed, developed and did maintenance of data integration programs in a Hadoop and RDBMS environment with both traditional and non-traditional source systems as we as RDBMS and NoSQL data stores for data access and analysis.
- Experienced in runningHadoopstreaming jobs to process terabytes of xml format data.
- Load and transform large sets of structured, semi structured and unstructured data.
- Strong experience of J2SE, XML, Web Services, WSDL, SOAP, UDDI, TCP, IP.
- Responsible to manage data coming from different sources.
- Assisted in exporting analyzed data to relational databases using Sqoop.
- Expert knowledge developing and debugging in Java/J2EE.
- Wrote Hive Queries and UDF’s.
- Developed Hive queries to process teh data and generate teh data cubes for visualizing.
- Extracted feeds form social media sites such as Facebook, Twitter.
- Created Pig Latin scripts to sort, group, join and filter teh enterprise wise data.
- Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
- Gained experience in managing and reviewing Hadoop log files.
Environment: Hadoop, Map Reduce, HDFS, Hive, Pig, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Java (jdk 1.6), Eclipse.
Confidential, Atlanta, GA
Hadoop Engineer
Responsibilities:
- Worked as ETL Architect to make sure all teh applications are migrated (along with server) smoothly.
- Deep understanding and related experience with Hadoop stack - internals, HBase, Hive, Pig and Map/Reduce.
- Wrote Hive Queries and UDF’s.
- Wrote MapReduce jobs.
- Upgrading teh Hadoop Cluster to CDH2 and setup High availability Cluster Integrate teh HIVE with existing applications.
- A deep and thorough understanding of ETL tools and how they can be applied in a Big Data environment.
- Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, manage and review data backups, manage and review Hadoop log files.
- Familiar with ETL Standards and Process and developed ETL logic as per standards from Source-Flat File, Flat-File-Stage, Stage-Work, Work-Work Interim tables and Work Interim tables- Target Tables
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
- Created Hive UDFS to extract data from staging tables.
- Extracted teh data from Teradata into HDFS using Sqoop.
- Analyzed teh data by performing Hive queries and running Pig scripts to know user behavior like shopping enthusiasts, travelers, music lovers etc.
- Continuous monitoring and managing teh Hadoop cluster through Cloudera Manager.
- Involved in moving all log files generated from various sources to HDFS for further processing through Flume.
- Created HBase tables to store variable data formats of data coming from different portfolios.
- Involved in transforming data from Mainframe tables to HDFS, and HBASE tables using Sqoop.
- Implemented test scripts to support test driven development and continuous integration.
- Specifying teh Cluster size, allocating Resource pool, Distribution of Hadoop by writing teh specification texts in JSON File format.
- Strong experience of software and system development using JSP, Servlet, Java Server Face, EJB, JDBC, JNDI, Struts, Maven, Trac, Subversion, JUnit, SQL language.
- Hands-on experience of Sun One Application Server, Web logic Application Server, Web Sphere Application Server, Web Sphere Portal Server, and J2EE application deployment technology.
- Responsible to manage data coming from different sources.
- Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
- Created and maintained Technical documentation for launching HADOOP Clusters and for executing Hive queries and Pig Scripts.
Environment: Hadoop, HDFS, MapReduce, Hive, Pig, Sqoop, Cygwin, Oracle, SQL Server. MySQL, UNIX Shell Scripting, SQL, PL/SQL, TOAD, Windows NT, SQL Server Management Studio.
Confidential, Phoenix, AZ
Hadoop Developer
Responsibilities:
- Involved in loading data from UNIX file system to HDFS.
- Installed and configured HadoopMapReduce, HDFS and developed multiple MapReduce jobs in Java for data cleansing and preprocessing.
- Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
- Devised procedures that solve complex business problems with due considerations for hardware/software capacity and limitations, operating times and desired results.
- Worked hands on with ETL process.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
- Expert knowledge developing and debugging in Java/J2EE.
- Developed application using Eclipse.
- Hands-on experience of Sun One Application Server, Web logic Application Server, Web Sphere Application Server, Web Sphere Portal Server, and J2EE application deployment technology.
- Created Hive tables to store teh processed results in a tabular format.
- Created HBase tables to store variable data formats of data coming from different portfolios.
- Created HBase tables to store variable data formats of data coming from different portfolios.
- Develop HIVE queries for teh analysts.
- Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
- Exported teh result set from HIVE to MySQL using Shell scripts.
- Used Zookeeper for various types of centralized configurations.
- Assembled human genome on 200 node cluster.
- Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, manage and review data backups, manage and review Hadoop log files.
- Continuous monitoring and managing teh Hadoop cluster using Cloudera Manager.
- Automated all teh jobs starting from pulling teh Data from different Data Sources like MySQL to pushing teh result set Data to Hadoop Distributed File System using Sqoop.
- Supported Map Reduce Programs those are running on teh cluster.
- Experienced in runningHadoopstreaming jobs to process terabytes of xml format data.
Environment: Hadoop, Map Reduce, HDFS, Hive, Ooozie, Zookeeper, Java (jdk1.6), Cloudera, NoSQL, Oracle 11g, 10g, PL SQL, SQL*PLUS, Toad 9.6, Windows NT, UNIX Shell Scripting.
Confidential - Windsor, CT
Java/J2ee developer
Responsibilities:
- Understanding teh Business Functionality and Application flow.
- Architected a JSF, Web sphere, Oracle, spring, and Hibernate based 24x7 Web application.
- Experience with creating and reviewing UI design specifications, developing prototypes and conducting usability tests
- Involved in Analysis, Design, developing sequence diagrams and class diagrams using UML notations and Developed Design documents for teh requirement specification.
- Used Springs Jdbc and DAO layers to offer abstraction for teh business from teh database related code (CRUD).
- Prepared Technical design document to display all W4 form based on state selection and Loading into ADP system.
- Designed and integrated custom code with teh application user interface.
- Developed teh presentation layer using CSS and HTMLtaken from Bootstrap to develop for multiple browsers including mobiles and tablets.
- Developed business layer using spring, Hibernate and DAOs.
- Develop and Format teh GUI and Form Layout.
- Develop, implement, and maintain an asynchronous, AJAX based rich client for improved customer experience.
- Developed interactive web pages using AJAX and JavaScript
- Used Struts Action class for displaying teh correct W4 form based on state selection.
- Used EJB Stateless Session Bean for writing all teh business logic.
- ConfiguredStruts tilesfor reusing view components as an application of J2EE composite pattern.
- Review teh code and deployed teh application in Web sphere application server.
- Used ANT for deployment, integration of teh application.
- Wrote J unit test cases for actions and classes to validate teh functionality.
- Wrote store procedure to fetch/insert data into Oracle 10g R2 database.
- Involved in deployment on different servers like Development and UAT/Preprod environment
Environment: Java, J2EE, HTML, CSS, JSP, JDBC, Struts, JavaBeans, EJB, JavaScript, XML, Ajax, Web sphere 5.1, SQL,PL/SQL, Toad, Oracle 10g R2, Windows & Unix, VSS, Test Director, J unit, ANT.
Confidential, Wood Land Hills, CA
J2EE Developer
Responsibilities:
- Involved in Requirements Gathering, Analysis, Development and Documentation.
- Development of Web application follows MVC pattern utilizing Farmers internal framework to implement teh Controller layer and to assist with rendering teh View.
- Created UML class diagrams that depict teh code’s design and its compliance with teh functional requirements.
- There will be single controller servlet which will handle all web requests.
- Performed usability testing for teh application using JUnit Test
- Development of front-end validations using JavaScript.
- Used Hibernate as teh ORM tool to communicate with teh database.
- Involved in writing session beans, message driven beans and hibernate mapping files and hibernate configuration files.
- Involved in Bug fixing and functionality enhancements.
- Developed Web pages using JSF, JSP and Ajax.
- Involved in unit testing, Integration activities and validation of teh application.
Environment: JDK 1.5, Log4j, JSP, Hibernate, Oracle 10g, JSP 2.0, XML, XSL, HTML, Java Script, JSF, and Servlets 2.4.
Confidential
Java Developer
Responsibilities:
- Technical responsibilities included high level architecture and rapid development.
- Design architecture following J2EE MVC framework.
- Developed interfaces using HTML, JSP pages and Struts -Presentation View.
- Involved in designing & developing web-services using SOAP and WSDL.
- Developed and implemented Servlets running under JBoss.
- Used J2EE design Patterns for teh Middle Tier development.
- Used J2EE design patterns and Data Access Object (DAO) for teh business tier and integration Tier layer of teh project.
- Created UML class diagrams that depict teh code’s design and its compliance with teh functional requirements.
- Developed various EJBs for handling business logic and data manipulations from database.
- Designed and developed teh UI using Struts view component, JSP, HTML, CSS and JavaScript.
- Implemented CMP entity beans for persistence of business logic implementation
- Development of database interaction code to JDBC API making extensive use of SQL Query Statements and advanced prepared statement.
- Involved in writing Spring Configuration XML files that contains declarations and other dependent objects declaration.
- Inspection/Review of quality deliverables such as Design Documents.
- Involved in creation running of Test Cases for JUnit Testing.
- Experience in implementing Web Services using SOAP, REST and XML/HTTP technologies.
- Used Log4J to print teh logging, debugging, warning, info on teh server console.
- Wrote SQL Scripts,Stored procedures and SQL Loader to load reference data.
Environment: J2EE (Java Servlets, JSP, Struts), MVC Framework, Apache Tomcat, JBoss, Oracle8i.
