We provide IT Staff Augmentation Services!

Etl - Hadoop Architect/developer Resume

3.00/5 (Submit Your Rating)

NJ

SUMMARY

  • Over 8 years of extensive software development experience in full life cycle development which includes more than 3 years of experience as Hadoop Developer/Architect focusing on various Big Data Technologies.
  • Experience in developing Map Reduce Programs using Apache Hadoop for analyzing teh big data as per teh requirement.
  • Hands on experience in writing MapReduce jobs in Java, Pig and Python.
  • Developing both batch and real time applications on teh Hadoop platform.
  • Experienced on major Hadoop ecosystem’s projects such as PIG, HIVE and HBASE.
  • Good working experience using Sqoop to import data into HDFS from RDBMS and vice - versa.
  • Good knowledge in using job scheduling and monitoring tools like Oozie and ZooKeeper.
  • Knowledge of NoSQL databases such as HBase, and MongoDB.
  • Good understanding/knowledge of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Has a good understanding to ETL concepts.
  • Hands on experience in installing, configuring and using ecosystem components like HadoopMapReduce, HDFS, Hbase, ZooKeeper, Oozie, Flume, Sqoop, Pig & Hive with CDH3&4&5 clusters..
  • Experience in managing Hadoop clusters using Cloudera Manager
  • Experience in database development using SQL and PL/SQL and experience working on databases like Oracle 9i/10g, Informix, and SQL Server.
  • Performeddata analysisusingMySQL, SQL Server Management Studio, and Oracle.
  • Profound knowledge of teh principal of DW using Fact Tables, Dimension Tables, Star schema modeling and Snowflake Schema modeling.
  • Strong skills in Datastage Administrator in UNIX & LINUX environments, Report creation using OLAP data source and having knowledge in OLAP universe.
  • Worked on debugging tools such as Dtrace, Struss and Top. Expert in setting up SSH, SCP, SFTP connectivity between UNIX hosts.
  • Experienced teh integration of various data sources like Java, RDBMS, Shell Scripting, Spreadsheets, and Text files.
  • Used Springs JDBC and DAO layers to offer abstraction for teh business from teh database related code (CRUD).
  • Working with relative ease with different working strategies like Agile, Waterfall and Scrum methodologies.
  • Expertise in designing and developing J2EE compliant systems using IDE tools like Eclipse, WebSphere Studio Application Developer (WSAD).
  • In-depth understanding of Data Structures and Algorithms.
  • UNIX shell scripting, resource Extensive experience in Java and J2EE technologies like Servlets, JSP, JSF, and JDBC.
  • Expert in using J2EE complaint application servers Apache Tomcat, IBM Web Sphere.
  • Extensively worked ondebuggingusing Eclipse debugger.
  • Hands on experience working with Java project build managers Apache MAVEN and ANT
  • Good Knowledge in Flume, Avro and Zoo Keeper Architecture.
  • Good understanding of Data Mining and Machine Learning techniques.
  • Extensive experience in working with different databases such as Oracle, IBM DB, RDBMS, SQL Server, MySQL and writing Stored Procedures, Functions, Joins and Triggers for different Data Models.
  • Extensive experience in working with Parallel jobs, Troubleshooting and Performance tuning.
  • Strong work ethic with desire to succeed and make significant contributions to teh organization.
  • Strong problem solving skills, good communication, interpersonal skills and a good team player.
  • Has teh motivation to take independent responsibility as well as ability to contribute and be a productive team member.

TECHNICAL SKILLS

Hadoop/Big Data Technologies: HDFS, Map Reduce, Hive, Pig, Sqoop, Flume, Oozie, and Zookeeper.

No SQL Databases: Hbase,Cassandra, mongoDB.

Languages: C, C++, Java, J2EE, C#, Asp.Net, PL/SQL, Pig Latin, HiveQL, Unix shell Scripts.

Java/J2EE Technologies: Applets, Swing, JDBC, JNDI, JSON, JSTL, RMI, JMS, Java Script, JSPServlets, EJB, JSF, JQuery.

Frameworks: MVC, Struts, Spring, And Hibernate.

Web Services: HTML, DHTML, XML, AJAX, WSDL, SOAP, REST, Jersey.

Web/Application servers: Apache Tomcat, WebLogic, JBoss, Web Sphere.

Operating Systems: Sun Solaris, HP-UNIX, RedHat Linux, Ubuntu Linux and Windows XP/Vista/7/8.

Databases: Oracle 9i/10g/11g, DB2, SQL Server, MySQL.

Web technologies: JSP, Servlets, JNDI, JDBC, Java Beans, JavaScript, Web Services (JAX-WS).

Java IDE: Eclipse 3.x, IBM Web Sphere Application Developer, IBM RAD 7.0.

Tools: TOAD, SQL Developer, SOAP UI, ANT, Maven, Visio, Rational RoseEndur 8.x/10.x/11.x.

PROFESSIONAL EXPERIENCE

Confidential, NJ

ETL - Hadoop Architect/Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Installed and configured Hive, Pig, Sqoop, Flume and Oozie on teh Hadoop cluster.
  • Developed Simple to complex Map/reduce Jobs using Hive and Pig.
  • Devised and lead teh implementation of teh next generation architecture for more efficient data ingestion and processing.
  • Give extensive presentations about teh Hadoop ecosystem, best practices, data architecture in Hadoop.
  • Provide mentorship and guidance to other architects to halp them become independent.
  • Provide review and feedback for existing physical architecture, data architecture and individual code.
  • Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Involved in loading data from UNIX file system to HDFS.
  • Wrote MapReduce jobs to discover trends in data usage by users.
  • Involved in running Hadoop streaming jobs to process terabytes of text data.
  • Analyzed large data sets by running Hive queries and Pig scripts.
  • Helped teh team to increase teh Cluster size from 22 to 30 Nodes.
  • Job management using Fair scheduler.
  • Develop Core Framework based on Hadoop to Migrate Existing ETL (RDBMS) Solution.
  • Wrote Pig Scripts to generate Map Reduce jobs and performed ETL procedures on teh data in HDFS
  • A deep and thorough understanding of ETL tools and how they can be applied in a Big Data environment.
  • Worked as ETL Architect to make sure all teh applications are migrated (along with server) smoothly.
  • Involved in creating Hive tables, and loading and analyzing data using hive queries.
  • Responsible for managing data from multiple source.
  • Designed, developed and did maintenance of data integration programs in a Hadoop and RDBMS environment with both traditional and non-traditional source systems as we as RDBMS and NoSQL data stores for data access and analysis.
  • Experienced in runningHadoopstreaming jobs to process terabytes of xml format data.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Strong experience of J2SE, XML, Web Services, WSDL, SOAP, UDDI, TCP, IP.
  • Responsible to manage data coming from different sources.
  • Assisted in exporting analyzed data to relational databases using Sqoop.
  • Expert knowledge developing and debugging in Java/J2EE.
  • Wrote Hive Queries and UDF’s.
  • Developed Hive queries to process teh data and generate teh data cubes for visualizing.
  • Extracted feeds form social media sites such as Facebook, Twitter.
  • Created Pig Latin scripts to sort, group, join and filter teh enterprise wise data.
  • Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
  • Gained experience in managing and reviewing Hadoop log files.

Environment: Hadoop, Map Reduce, HDFS, Hive, Pig, Java, SQL, Cloudera Manager, Sqoop, Flume, Oozie, Java (jdk 1.6), Eclipse.

Confidential, Atlanta, GA

Hadoop Engineer

Responsibilities:

  • Worked as ETL Architect to make sure all teh applications are migrated (along with server) smoothly.
  • Deep understanding and related experience with Hadoop stack - internals, HBase, Hive, Pig and Map/Reduce.
  • Wrote Hive Queries and UDF’s.
  • Wrote MapReduce jobs.
  • Upgrading teh Hadoop Cluster to CDH2 and setup High availability Cluster Integrate teh HIVE with existing applications.
  • A deep and thorough understanding of ETL tools and how they can be applied in a Big Data environment.
  • Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, manage and review data backups, manage and review Hadoop log files.
  • Familiar with ETL Standards and Process and developed ETL logic as per standards from Source-Flat File, Flat-File-Stage, Stage-Work, Work-Work Interim tables and Work Interim tables- Target Tables
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
  • Created Hive UDFS to extract data from staging tables.
  • Extracted teh data from Teradata into HDFS using Sqoop.
  • Analyzed teh data by performing Hive queries and running Pig scripts to know user behavior like shopping enthusiasts, travelers, music lovers etc.
  • Continuous monitoring and managing teh Hadoop cluster through Cloudera Manager.
  • Involved in moving all log files generated from various sources to HDFS for further processing through Flume.
  • Created HBase tables to store variable data formats of data coming from different portfolios.
  • Involved in transforming data from Mainframe tables to HDFS, and HBASE tables using Sqoop.
  • Implemented test scripts to support test driven development and continuous integration.
  • Specifying teh Cluster size, allocating Resource pool, Distribution of Hadoop by writing teh specification texts in JSON File format.
  • Strong experience of software and system development using JSP, Servlet, Java Server Face, EJB, JDBC, JNDI, Struts, Maven, Trac, Subversion, JUnit, SQL language.
  • Hands-on experience of Sun One Application Server, Web logic Application Server, Web Sphere Application Server, Web Sphere Portal Server, and J2EE application deployment technology.
  • Responsible to manage data coming from different sources.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and pig jobs.
  • Created and maintained Technical documentation for launching HADOOP Clusters and for executing Hive queries and Pig Scripts.

Environment: Hadoop, HDFS, MapReduce, Hive, Pig, Sqoop, Cygwin, Oracle, SQL Server. MySQL, UNIX Shell Scripting, SQL, PL/SQL, TOAD, Windows NT, SQL Server Management Studio.

Confidential, Phoenix, AZ

Hadoop Developer

Responsibilities:

  • Involved in loading data from UNIX file system to HDFS.
  • Installed and configured HadoopMapReduce, HDFS and developed multiple MapReduce jobs in Java for data cleansing and preprocessing.
  • Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
  • Devised procedures that solve complex business problems with due considerations for hardware/software capacity and limitations, operating times and desired results.
  • Worked hands on with ETL process.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
  • Expert knowledge developing and debugging in Java/J2EE.
  • Developed application using Eclipse.
  • Hands-on experience of Sun One Application Server, Web logic Application Server, Web Sphere Application Server, Web Sphere Portal Server, and J2EE application deployment technology.
  • Created Hive tables to store teh processed results in a tabular format.
  • Created HBase tables to store variable data formats of data coming from different portfolios.
  • Created HBase tables to store variable data formats of data coming from different portfolios.
  • Develop HIVE queries for teh analysts.
  • Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
  • Exported teh result set from HIVE to MySQL using Shell scripts.
  • Used Zookeeper for various types of centralized configurations.
  • Assembled human genome on 200 node cluster.
  • Responsible for cluster maintenance, adding and removing cluster nodes, cluster monitoring and troubleshooting, manage and review data backups, manage and review Hadoop log files.
  • Continuous monitoring and managing teh Hadoop cluster using Cloudera Manager.
  • Automated all teh jobs starting from pulling teh Data from different Data Sources like MySQL to pushing teh result set Data to Hadoop Distributed File System using Sqoop.
  • Supported Map Reduce Programs those are running on teh cluster.
  • Experienced in runningHadoopstreaming jobs to process terabytes of xml format data.

Environment: Hadoop, Map Reduce, HDFS, Hive, Ooozie, Zookeeper, Java (jdk1.6), Cloudera, NoSQL, Oracle 11g, 10g, PL SQL, SQL*PLUS, Toad 9.6, Windows NT, UNIX Shell Scripting.

Confidential - Windsor, CT

Java/J2ee developer

Responsibilities:

  • Understanding teh Business Functionality and Application flow.
  • Architected a JSF, Web sphere, Oracle, spring, and Hibernate based 24x7 Web application.
  • Experience with creating and reviewing UI design specifications, developing prototypes and conducting usability tests
  • Involved in Analysis, Design, developing sequence diagrams and class diagrams using UML notations and Developed Design documents for teh requirement specification.
  • Used Springs Jdbc and DAO layers to offer abstraction for teh business from teh database related code (CRUD).
  • Prepared Technical design document to display all W4 form based on state selection and Loading into ADP system.
  • Designed and integrated custom code with teh application user interface.
  • Developed teh presentation layer using CSS and HTMLtaken from Bootstrap to develop for multiple browsers including mobiles and tablets.
  • Developed business layer using spring, Hibernate and DAOs.
  • Develop and Format teh GUI and Form Layout.
  • Develop, implement, and maintain an asynchronous, AJAX based rich client for improved customer experience.
  • Developed interactive web pages using AJAX and JavaScript
  • Used Struts Action class for displaying teh correct W4 form based on state selection.
  • Used EJB Stateless Session Bean for writing all teh business logic.
  • ConfiguredStruts tilesfor reusing view components as an application of J2EE composite pattern.
  • Review teh code and deployed teh application in Web sphere application server.
  • Used ANT for deployment, integration of teh application.
  • Wrote J unit test cases for actions and classes to validate teh functionality.
  • Wrote store procedure to fetch/insert data into Oracle 10g R2 database.
  • Involved in deployment on different servers like Development and UAT/Preprod environment

Environment: Java, J2EE, HTML, CSS, JSP, JDBC, Struts, JavaBeans, EJB, JavaScript, XML, Ajax, Web sphere 5.1, SQL,PL/SQL, Toad, Oracle 10g R2, Windows & Unix, VSS, Test Director, J unit, ANT.

Confidential, Wood Land Hills, CA

J2EE Developer

Responsibilities:

  • Involved in Requirements Gathering, Analysis, Development and Documentation.
  • Development of Web application follows MVC pattern utilizing Farmers internal framework to implement teh Controller layer and to assist with rendering teh View.
  • Created UML class diagrams that depict teh code’s design and its compliance with teh functional requirements.
  • There will be single controller servlet which will handle all web requests.
  • Performed usability testing for teh application using JUnit Test
  • Development of front-end validations using JavaScript.
  • Used Hibernate as teh ORM tool to communicate with teh database.
  • Involved in writing session beans, message driven beans and hibernate mapping files and hibernate configuration files.
  • Involved in Bug fixing and functionality enhancements.
  • Developed Web pages using JSF, JSP and Ajax.
  • Involved in unit testing, Integration activities and validation of teh application.

Environment: JDK 1.5, Log4j, JSP, Hibernate, Oracle 10g, JSP 2.0, XML, XSL, HTML, Java Script, JSF, and Servlets 2.4.

Confidential

Java Developer

Responsibilities:

  • Technical responsibilities included high level architecture and rapid development.
  • Design architecture following J2EE MVC framework.
  • Developed interfaces using HTML, JSP pages and Struts -Presentation View.
  • Involved in designing & developing web-services using SOAP and WSDL.
  • Developed and implemented Servlets running under JBoss.
  • Used J2EE design Patterns for teh Middle Tier development.
  • Used J2EE design patterns and Data Access Object (DAO) for teh business tier and integration Tier layer of teh project.
  • Created UML class diagrams that depict teh code’s design and its compliance with teh functional requirements.
  • Developed various EJBs for handling business logic and data manipulations from database.
  • Designed and developed teh UI using Struts view component, JSP, HTML, CSS and JavaScript.
  • Implemented CMP entity beans for persistence of business logic implementation
  • Development of database interaction code to JDBC API making extensive use of SQL Query Statements and advanced prepared statement.
  • Involved in writing Spring Configuration XML files that contains declarations and other dependent objects declaration.
  • Inspection/Review of quality deliverables such as Design Documents.
  • Involved in creation running of Test Cases for JUnit Testing.
  • Experience in implementing Web Services using SOAP, REST and XML/HTTP technologies.
  • Used Log4J to print teh logging, debugging, warning, info on teh server console.
  • Wrote SQL Scripts,Stored procedures and SQL Loader to load reference data.

Environment: J2EE (Java Servlets, JSP, Struts), MVC Framework, Apache Tomcat, JBoss, Oracle8i.

We'd love your feedback!