Sr. Hadoop Developer Resume
Menlo Park, CA
SUMMARY:
- More than 7 years of experience in Hadoop/Big Data technologies such as Hadoop, Pig, Hive, HBase, Oozie, Zookeeper, Sqoop, Storm, Flink, Flume, Zookeeper, Impala, Tez, Kafka Spark and Map Reduce/YARN.
- Expertise in Tableau BI reporting tools & Tableau Dashboards Developments & Server Administration.
- Hands - on experience in using Hive partitioning, bucketing and execute different types of joins on Hive tables and implementing Hive SerDes like JSON and Avro.
- Experience in Batch and Real Time Processing data on the Hadoop Ecosystem.
- Expertise in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, Sqoop, Flume, Spark, Impala.
- Expertise in different data loading techniques (Flume, Sqoop, Spark) onto HDFS.
- Exceptional skills with NoSQL databases such as HBase and Cassandra.
- Strong knowledge in data modelling, effort estimation, ETL Design, development, system testing, implementation and production support. Experience in resolving on-going maintenance issues and bug fixes.
- Good knowledge of Normalization, Fact Tables and Dimension Tables, also dealing with OLAP and OLTP systems.
- Strong experience in designing and developing various Data Mart and Data Warehouse applications using custom ETL Frameworks developed in Java and Python.
- Experience in data ingress and egress using Sqoop from HDFS to Relational Database Systems and vice-versa.
- Hands on experience in writing Pig Latin Scripts, Hive-QL queries, Spark SQL.
- Expert knowledge in real time data analytics using Apache Storm.
- Expertise in Java/J2EE technologies such as Core Java, spring, Hibernate, JDBC, JSON, HTML, Struts, Servlets, JSP, JBOSS and JavaScript.
- Experience of developing SQL scripts using Spark for handling different data sets and verifying the performance over Map Reduce jobs.
- Experience in migrating the data using Sqoop from Hadoop to Relational Database System and vice-versa.
- Worked in Agile methodology of software development process as a Scrum Master.
- Strong foundation in Programming, debugging skills, developed modules which have met with client requirements & targets.
- Expert in creating UDF’s, UDTF’s and UDAF’s for Hive, Pig and Impala.
- Expertise in Hadoop administration such as managing cluster, reviewing Hadoop log files.
- Good Experience in writing complex SQL queries with databases like DB2, Oracle 10g, MySQL, SQL Server and MS SQL Server 2005/2008.
- Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.
TECHNICAL SKILLS:
- Expertise in Business tools like Tableau 8.X/9.X, Micro-Strategy, Business Objects XI R2, Informatica Power center, OLAP/OLTP, Dimension Modeling, Data Modeling, Microsoft SQL Server, PHP, Pentaho, Talend.
- Experience in Big data tools like Hadoop, Cloudera Manager, Map Reduce 1.0/2.0, Pig, Hive, HBase, Sqoop, Oozie, Zookeeper, Avro, Kafka, Spark, Flume, Storm, Impala, Scala, Mahout, Hue.
- Good Knowledge of Databases like DB2, MySQL, MS Access, MS SQL server, Teradata, Vertica, Aster n-Cluster, SSAS, Oracle, Presto-DB.
- Experience in languages like Java / J2EE, Scala, Python HTML, SQL, spring, Hibernate, JDBC, JSON, JavaScript, Spark SQL.
- Experience in Operating systems like Mac OS, Unix, Linux (Various Versions), Windows 2003/7/8/8.1/10/ XP/Vista
- Expertise in Web Application Servers like Apache Tomcat, WebLogic, WebSphere Tools Eclipse, NetBeans.
- Good Knowledge of Version tools like Git, SVN, Perforce and Mercury.
PROFESSIONAL EXPERIENCE:
Sr. Hadoop Developer
Confidential, Menlo Park, CA
Responsibilities:
- Development and testing of Hadoop jobs and implemented data quality solution based on design.
- Developed data pipeline using Flume to ingest customer behavioral data into HDFS for analysis.
- Hands-on experience in using Hive partitioning, bucketing and execute different types of joins on Hive tables and implementing Hive SerDes like JSON and Avro.
- Written Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data.
- Involved in developing Pig Scripts for change data capture and delta record processing between newly arrived data and already existing data in HDFS.
- In data exploration stage used hive and impala to get some insights about the customer data.
- Used Hive data warehouse tool to analyze the data in HDFS and developed Hive queries.
- Created HBase tables to store variable data formats of PII data coming from different portfolios.
- Developed interactive Dashboards and reports using Tableau server.
- Developed multiple test cases for calculated fields and developed SQL
- Created action filters, parameters and calculated sets for preparing dashboards and worksheets in Tableau.
- Exported data from Impala to Tableau reporting tool, created dashboards on live connection.
- Consolidating data from various heterogeneous systems into unified format to be consumed by Client Profitability report.
- Created dynamic BI report/dashboard for production support in Excel/PowerPoint/Power BI/Tableau/ My SQL Server/ PHP.
- Designed Data Model in Power BI such as Fact and dimensional tables and their relationships
- Used Informatica to extract, transform & load data from SQL Server to Oracle databases.
- Worked on Micro Strategy mobile platform for mobile application development for the Sales Portal App within Confidential .
- Experience in Migrating objects between different environment Dev, Test and Prod.
Sr. Hadoop Developer
Confidential, Denver, Colorado
Responsibilities:
- Developed a data pipeline using Kafka, Spark and Hive to ingest, transform and analyzing data.
- Development and testing of ETL jobs and implemented data quality solution based on design.
- Used Pig as ETL tool to do Transformations, even joins and some pre-aggregations before storing the data on to HDFS.
- Developed HIVE scripts for analyst requirements for analysis. Cross-examining data loaded in Hive table with the source data in oracle.
- Used IMPALA to analyze data ingested into HBase and compute various metrics for reporting on the dashboard.
- Worked on data migration and data conversion using PL/SQL, SQL and Python to convert them into custom ETL tasks.
- Involved in creating Hive internal and external tables, loading them with data and writing hive queries which requires multiple join scenarios. Created partitioned and bucketed tables in Hive based on the hierarchy of the dataset.
- Worked on performance tuning of ETL code and SQL queries to improve performance, availability and throughput.
- Involved in ingesting data into Data Warehouse using various data loading techniques.
- Importing and Exporting of data from RDBMS to HDFS and vice versa using Sqoop
- Involved in preparing Best practices document, Code review methodology document, Migration document.
- Tuned existing ETL SQL scripts by making necessary design changes to improve performance of Fact tables.
- Wrote Scripts to generate Map Reduce jobs and performed ETL procedures on the data in HDFS using PIG, Hive, Python and OOZIE.
- Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
Sr. Hadoop Developer
Confidential, Chicago, Illinois
Responsibilities:
- Developed data pipelines using Sqoop, Flume, Pig and MapReduce to ingest historical and behavioral data into the Hadoop file systems to perform analysis.
- Designed Views within SQL Server 2012 to support Micro strategy reporting and dashboards.
- Loading of the data from various sources to Teradata using Pentaho.
- Collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis.
- Writing complex SQL Queries, Stored procedures, Triggers, Views, DDL, DML and UDFs to implement the business logic.
- Worked with End Users to identify the Report and Dashboard Requirements.
- Developed the PIG code for loading, filtering and storing the data.
- Developed a real-time analytics project using Kafka and Spark in Scala for optimizing ad placements on Confidential .com website.
- Worked on different levels of ETL loads like Extract Data from Source system to staging and then to load data into Warehouse.
- Worked extensively in tuning the current ETL processes for improving the performance by implementing database partitioning and increasing block size, data cache size and SQL overrides
- Developed Oozie workflow for scheduling and orchestrating the ETL process.
- Involved in creating Hive tables, loading data, and writing Hive queries.
- Developed data pipeline using Flume, Sqoop, Pig and Java map reduce to ingest customer behavioral data and financial histories into HDFS for analysis.
Sr. ETL Developer/ Data warehouse Consultant
Confidential, Dallas, Texas
Responsibilities:
- Responsible for loading the customer's data and event logs from Kafka into HBase using REST API.
- Worked on Tableau with crosstab functionality, Power pivot and experienced with Power BI reports and Power BI desktop
- Worked on debugging, performance tuning and Analyzing data using Hadoop components Hive & Pig.
- Created Hive tables from JSON data using data serialization framework like AVRO.
- Extensively involved in upgrading/migrating various ETL jobs from one version to another.
- Implemented generic export framework for moving data from HDFS to RDBMS and vice-versa.
- Worked on loading data from LINUX file system to HDFS.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Responsible for processing unstructured data using Pig and Hive.
- Worked on different levels of ETL loads like Extract Data from Source system to staging and then to load data into Warehouse.
- Adding nodes into the clusters & decommission nodes for maintenance.
- Extensive experience in managing and reviewing Hadoop log files.
- Very good understanding of Partitions, Bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
- Worked on various Business Object Reporting functionalities such as Slice and Dice, Master/detail, User Response function and different Formulas.
ETL Developer/Data warehouse Consultant
Confidential, Palo Alto, CA
Responsibilities:
- Involved in the design and development of Data Warehouse.
- Extensively used SQL and PL/SQL for development of Procedures, Functions, Packages and Triggers.
- Involved in deploying, configuring and managing reports using Report Manager and Report Builder
- Developed and Supported Map Reduce Programs those are running on the cluster.
- Extensively used Informatica Power Center 9.5/8.6.1 as ETL tool for developing the project.
- Interacted with business users on regular basis to consolidate and analyze the requirements and presented them with design results.
- Involved in data visualization and provided the files required for the team by analyzing the data in hive and impala for advanced analytics on the data
- Created many user-defined routines, functions, before/after subroutines which facilitated in implementing some of the complex logical solutions.
- Worked on improving the performance by using various performance tuning strategies.
- Managed the evaluation of ETL and OLAP tools and recommended the most suitable solutions depending on business needs.
- Migrated jobs from development to test and production environments.
Data Engineer
Confidential
Responsibilities:
- Involved in Requirements analysis, design, and development and testing.
- Involved in loading the created HFiles into HBase for faster access of large customer base without taking Performance hit.
- Used JDBC, SQL and PL/SQL programming for storing, retrieving, manipulating the data.
- Worked on client requirement and wrote Complex SQL Queries to generate Crystal Reports
- Extracted data from various sources using transformations and populated the data marts.
- Involved in creating Hive tables loading data and writing queries that will run internally in MapReduce way.
- Created SSIS packages to import data from MS Access, Excel to SQL 2008R2
- Scheduled the monthly/weekly/daily reports to run automatically.
- Use DDL and DML for writing triggers, stored procedures, and data manipulation
- Enhanced the currently built dimensional model using Star schema and implemented SCD that will help to track the history .
- Used SQL Agent to automatic process of extract, load data from source tables.
- Coded Java Servlets to control and maintain the session state and handle user requests.
- Worked on Data Partitioning, Snapshot Isolation in SQL Server 2008
- Responsible for gathering business requirements, developing and reviewing of design documents
- Wrote complex SQL queries and stored procedures.
- Used Maven to build the J2EE application.
- Involved in maintenance of different applications.
Environment: s: ETL, Hadoop, HDFS, HBase, Pig, Hive, Spark, Presto, Data-Swarm, Python, Vertica, Atom, Impala, Cloudera Manager, Informatica Power Center, Scala, Oracle, SQL server, SQL, PL/SQL, UNIX Scripting, Java, Teradata, Zookeeper, Oozie, JDK, Hue, Yarn, Flume, Kafka, Tableau etc.
