Sr. Hadoop Developer Resume
Atlanta, GA
PROFESSIONAL SUMMARY:
- 7+ years’ of Professional experience in IT Industry in Developing, Implementing, configuring, testing Hadoop ecosystem components and maintenance of various web based applications using Java, J2EE.
- 4+ years’ Real time experience in Hadoop Framework and its ecosystem.
- Experience in installation, configuration and managing - Cloudera’s Hadoop platform along with CDH3&4 clusters.
- Excellent knowledge on Hadoop Architecture and ecosystems such as HDFS, MapReduce, Job Tracker, Task Tracker, Namenode, Datanode and Secondary Namenode concepts.
- Knowledge in installing, configuring, and using Hadoop ecosystem components like Map Reduce, HDFS, Oozie, Hive, Sqoop, Pig, Flume and Zookeeper.
- Experience working on Sqoop incremental data loads and implementing Hive best practices like Partitioning, Bucketing and compressions.
- Knowledge on Spark and its in-memory capabilities, mainly in framework exploration for transition from MapReduce to Spark. Implemented Spark RDDS and Dataframes.
- Implemented PySpark using Jupyter notebooks and used SparkML library, Python NLTK to implement various machine learning algorithms.
- Experience using NoSQL database: HBase and its column-oriented storage.
- Expertise in writing Hadoop Jobs for processing and analyzing data using MapReduce, Hive & Pig. Experienced in extending Hive and Pig core functionality by writing custom UDFs using Java.
- Hands-on experience on YARN (MapReduce 2.0) architecture and it components.
- Excellent knowledge on implementing web services using REST and SOAP.
- Hands on experience in application development using Core Java, UNIX Shell scripting and RDBMS.
- Vast Experience in SQL Server and involved in migrating Projects from SQL Server to Hadoop.
- Exceptional skills in Agile Development using GitHub/Bitbucket.
- Excellent interpersonal and communication skills, result-oriented with problem solving and leadership skills.
TECHNICAL SKILLS:
Hadoop/Big: DataHDFS, MapReduce, Hive, Pig, Oozie, Sqoop, Spark, HBase, Zookeeper, YARN, TEZ, Flume
Java & J2EE Technologies: Core Java, Servlets, JSP, JDBC, JNDI and Java Beans
Frameworks: MVC, Struts, Hibernate, Spring
DatabasesS: QLServer, MySQL, DB2, SQL, NoSQL (Hbase)
Web Technologies: JavaScript, SOAP, AJAX, HTML, XML and CSS.
Programming Languages: SQL, Java, Python, UNIX Shell Scripting
IDE’Eclipse, Net: Beans, Jupyter Notebooks
Web Servers: Web Logic, Web Sphere, Apache Tomcat 6
Build Management tools: Maven, Apache ANT
Visualization Tools: Tableau, Arcadia Data, Apache Zeppelin
PROFESSIONAL EXPERIENCE:
Confidential, Atlanta, GA
Sr. Hadoop Developer
Responsibilities:
- Maintained client’s data in Test and Production Hadoop clusters.
- Imported client’s data from various data sources and staged it in the Hadoop datalake using Apache Sqoop.
- Worked along with data scientists and built huge de-normalized reporting Hive tables as required.
- SQL, streaming and complex analytics in the company are handled with use of PySpark. Used Eclipse with python plugins and Jupyter notebooks extensively to build PySpark codes.
- Used SparkSQL to read and write Hive tables, scripted complex hive queries on billions of records and executed them through SparkSQL to use the in-memory capabilities of Spark engine, thus achieving speed of execution.
- Worked with Spark RDDs and Dataframes to perform text cleansing on various text fields.
- Implemented several machine learning algorithms in Spark MLLib and Python NLTK.
- Submitted spark jobs in yarn-client and Yarn-cluster modes and implemented dynamic resource allocation.
- Integrated Spark with HBase for storing the processed log data into separate column families.
- Used complex data types like bags, tuples, and maps in PIG for handling data.
- Worked with Flume to import the log data from the reaper logs and syslog’s into the Hadoop cluster.
- Processed data into HDFS by developing solutions, analyzed the batch data using Map Reduce, Pig, Hive and produce summary results fromHadoopto downstream systems.
- Involved in creating Hive tables, Hive UDFs for missing functionality as per the business requirements for analytics.
- Implemented performance optimizations like using distributed cache for small datasets, Partitioning &Bucketing.
- Used different file formats like Text files, Sequence Files, Avro, Orc File, and RCF.
- Used various Compression Techniques like gzip, bzip2, snappy, lzo.
- Used SQOOP to dump data from Oracle and SQLServer into HDFS using incremental load. Involved actively verifying and testing data in HDFS and Hive tables while Sqooping data from RDBMS tables to Hive.
- Used Oozie to chain multiple Sqoop, Hive and Pig jobs which run independently with time and data availability.
- Secure the data imported into Hadoop datalake using standard security mechanisms and tools like Apache Ranger.
- Perform incremental data load into Hadoop datalake based on the input data frequency.
- Used CronTab and Cisco Tidal enterprise scheduler to schedule several jobs in production.
- Used Tableau and Arcadia Data to visualize the results of several analytics for better insights to business users.
- Worked in Agile environment using Git and Bitbucket.
Environment: Hortonworks HDP 2.6, Spark 1.6, HDFS, Pig, Sqoop, PySpark, Jupyter Notebooks, MapReduce, TEZ, HIVE, PIG, NoSQL - HBase, Zookeeper, Shell Scripting, Ubuntu, Flume, Tableau, Agile, Impala, Putty.
Confidential, Englewood, Colorado
Hadoop Developer
Responsibilities:
- Installed and configured Hadoop MapReduce, HDFS, Developed multiple MapReduce jobs in Java for data cleansing and processing.
- Written MapReduce code to process and parsing the data from various sources and storing the parsed data into HBase and Hive using Hive integration.
- Worked with Hbase and Hive scripts to extract, transform and load the data into HBase and Hive.
- Worked with different File Formats like TEXTFILE, AVROFILE, and ORC for HIVE querying and processing.
- Used various Compression Techniques like gzip, bzip2, snappy, lzo.
- Worked on installing cluster, commissioning & decommissioning of Datanode, Namenode recovery, capacity planning and slots configuration.
- Implemented Flume to collect data from various sources and is loaded into HDFS for further processing.
- Developed workflows using custom MapReduce, Pig, Hive and Sqoop.
- Tuned the cluster for optional performance to process these large datasets.
- Built reusable Hive UDFs to sort structure fields and return complex datatype.
- Exposure to performance tuning and optimization of Hive Queries.
- Responsible for loading data from UNIX file system to HDFS.
- Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using MR testing library.
- Developed workflow in Control-M to automate tasks of loading data into HDFS, preprocessing with PIG.
- Used Maven extensively for building jar files of MapReduce programs and deployed to cluster.
- Implemented Oozie engine to chain multiple MapReduce, Hive jobs.
- Modelled Hive Partitions extensively for data separation and faster data processing and followed Pig and Hive best practices for tuning.
- Partitioned each day’s data into separate partitions for easy access and efficiency.
- Used Sqoop to import from different databases like MySql, SQLServer and file systems to HDFS and vice versa.
- Exported the analyzed data to relational databases using Sqoop for visualization and to generate reports.
- Integrated BI tool with Impala.
- Cluster coordination services through Zookeeper.
- Also assisted admin team in installed and configuration of additional nodes in Hadoop cluster
- Used Visualization tools such as Power view for excel, Tableau for visualizing and generating reports.
Environment: Hadoop, MapReduce, Hive, MySQL, Hbase, HIVE Impala, PIG, Sqoop, Oozie, Flume, Cloudera, Zookeeper, Eclipse (Kepler), SQLServer, MySql, Toad 9.6, UNIX, Tableau, Control-M
Confidential, Cincinnati, OH
Hadoop Developer
Responsibilities:
- Interacted daily with the onshore counterparts to gather requirements.
- Worked on few development projects on Microsoft SQLServer and gained knowledge on SSIS and SSAS.
- Performed Input data analysis, generated space estimation reports for Staging and Target tables in Testing and Production environments.
- Converted projects from SQLServer to Hadoop to handle the increasing data of the Bank and make use of its distributed file system and parallel processing capabilities, thus gained knowledge on Hadoop right from scratch.
- Installed and configured Hadoop MapReduce, HDFS and started to load data into HDFS instead of SQLServer and performed Data Cleansing and Processing operations.
- Have gained an in-depth knowledge on HDFS data storage and MapReduce data processing techniques.
- Performed importing and exporting data into HDFS and Hive using Sqoop.
- Designed and Scheduled workflows using Oozie.
- Created UNIX Shell Scripts to with in-turn call several Hive and Oozie scripts.
- Built several Managed and External HIVE tables and performed several Joins on those tables to achieve the result.
- Modelled Hive Partitions extensively for data separation and faster data processing and followed Pig and Hive best practices for tuning.
- Implemented Static and Dynamic Partitioning in Hive. Partitioned data into separate partitions based on year, gender, dates etc. for easy access and efficiency.
- Worked on a Remediation project to optimize all the SQL queries in e-commerce line of business and applied several query tuning and query optimization techniques in SQL.
- Scheduled several jobs which include several Shell, Oozie, hive and SQOOP scripts using Autosys. Coded JIL scripts to determine the job dependencies while scheduling.
- Performed Data analytics on the credit card transactional data and Coded Automatic Report Mailing Scripts in UNIX.
- Worked in Testing and Production environments and learnt moving components from one environment to the other Subversion (SVN).
- Performed Encryption and Decryption of key fields like Account Number of the Bank customers in several input and reporting files using COBOL.
- Built complete ETL logic, generated transformations, work flows and automated the scheduled runs.
- Built an Automatic Query Performance Metrics Generation Tool in SQLServer using SQL Procedures.
- Generated Test scripts and Test plan, Data Analysis and Defect Reporting using HP Quality Center.
Environment: Hadoop, SQLServer, HDFS, MapReduce, HIVE, PIG, Sqoop, Oozie, Java, COBOL, Autosys, HP Quality Center, SVN.
Confidential, New York, NY
Java Developer
Responsibilities:
- Used Hibernate ORM tool as persistence Layer - using the database and configuration data to provide persistence services (and persistent objects) to the application.
- Implemented Oracle Advanced Queuing using JMS and Message driven beans.
- Responsible for developing DAO layer using Spring MVC and configuration XML’s for Hibernate and to also manage CRUD operations (insert, update, and delete).
- Implemented Dependency injection of Spring frame work.
- Developed and implemented the DAO and service classes.
- Developed reusable services using BPEL to transfer data.
- Participated in Analysis, interface design and development of JSP.
- Configured log4j to enable/disable logging in application.
- Wrote SPA (Single Page Web Applications) using RESTFUL web services plus Ajax and AngularJS.
- Developed Rich user interface using HTML, JSP, AJAX, JSTL, Java Script, JQuery and CSS.
- Implemented PL/SQL queries, Procedures to perform data base operations.
- Wrote UNIX Shell scripts and used UNIX environment to deploy the EAR and read the logs.
- Implemented Log4j for logging purpose in the application.
- Involved in code deployment activities for different environments.
- Implemented agile development methodology.
Environment: Java, Spring, Hibernate, JMS, EJB, Web logic Server, JDeveloper, SQL Developer, Maven, XML, CSS, JSON, JavaScript.
Confidential
Java Developer
Responsibilities:
- Developed the user interface screens using web technologies like JSP, JavaScript, HTML, AJAX, CSS, and XSLT.
- Modeling conceptual design using Use Case, UML Class and Activity diagrams using Rational Rose.
- Wrote requirement specific SQL and PL/SQL scripts including Stored Procedures, functions, packages and triggers.
- Implemented Database access through JDBC at Server end with Oracle.
- Used Spring Aspect Oriented Programming (AOP) for addressing cross cutting concerns.
- Developed request/response paradigm by using Spring Controllers, Inversion of Control and Dependency Injection with Spring MVC.
- Used Web Services like SOAP and WSDL to communicate over internet.
- Involved in implementation of the JMS Connection Pool, including publish and subscribe using Spring JMS.
- Used CVS for version control and Log4j for logging.
- Used JProbe and JConsole to profile application for memory leaks and resource utilization.
- Developed test classes in JUnit for implementing unit testing.
- Deployed the application using WebLogic Application Server.
- Made enhancements to the application which presented me with the opportunity to go through the entire SDLC.
- Providing daily updates to the on-site team over call and making enhancements.
- Assisted the Quality Assurance team in testing the applications.
- Code coverage and Test case presentation.
