We provide IT Staff Augmentation Services!

Etl Developer Resume

2.00/5 (Submit Your Rating)

Irving, TX

SUMMARY:

  • 7 years of professional IT experience including 4 years hands - on experience in Hadoop ecosystem components.
  • Expertise in providing solutions for Big Data using Hadoop 2.x, HDFS, MR2, YARN, Kafka, PIG, Hive, Spark, SQOOP, HBase, Cloudera Manager, Zoo keeper, Oozie, Hue.
  • Experience in Cloudera and Hortonworks distributions on Hadoop.
  • Experienced configuring various Hadoop components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce2, and YARN.
  • Experienced in using Kafka as a distributed publisher-subscriber messaging system.
  • Experience in importing and exporting data using SQOOP from HDFS/Hive/HBase to Relational Database Systems and vice-versa.
  • Hands on experience in in-memory data processing with Apache Spark applications utilizing dataframe and spark SQL API.
  • Good experience in writing PIG scripts and Hive Queries for processing and analyzing large volumes of data.
  • Experience in Partitioning, Bucketing, Join Optimizations and query optimizations in Hive and automating the Hive Queries with dynamic partitioning.
  • Experience in optimization of MapReduce algorithm using Combiners and Partitioners to deliver best results.
  • Expertise in extending Hive and Pig core functionality by writing custom UDFs to meet certain business requirements.
  • Proficient in designing and ingesting data to NoSQL database like HBase.
  • Experience in managing and reviewing Hadoop log files.
  • Highly motivated and versatile team player with the ability to work independently & adapt quickly to new emerging technologies.

TECHNICAL SKILLS:

Big Data Technologies: HDFS, MapReduce, Hive, Spark, HBase, Pig, SQOOP, Oozie, Zookeeper, Flume, Kafka

Scripting Languages: Shell Scripting, Unix Script, Python, Scala

Programming Languages: Python, Java, C, SQL, HQL

RDBMS: MySQL 5.5, MSSQL 2014, Oracle 10G

IDE s: NetBeans, Eclipse, Microsoft Visual Studio, IntelliJ IDEA

Virtual Machines: VMWare, Virtual Box

Operating Systems: Cent OS 5.5, Unix, Red Hat Linux, Ubuntu, Windows

Web Technologies: HTML, CSS, JavaScript, Apache Tomcat

File Formats: XML, Text, Sequence, RC, JSON, ORC, AVRO, and Parquet etc.

PROFESSIONAL EXPERIENCE:

Confidential, Irving, TX

HADOOP DEVELOPER

Responsibilities:

  • Involved in managing nodes on Hadoop cluster and monitor Hadoop cluster job performance using Cloudera manager.
  • Performed analytics in Hive using various files format like JSON, Avro, ORC, and Parquet.
  • Worked on Spark RDD transformations to map business analysis and apply actions on top of transformations.
  • Experienced in performing analytics using Spark SQL and Scala on different formats like Text file, Avro, Parquet files.
  • Worked on Spark streaming to get ongoing information from Kafka and store the stream information to HDFS.
  • Developed Pig Latin scripts and Pig command line transformations for data joins and custom processing of Map reduce outputs.
  • Automated jobs for data ingestion, enrichment, and provisioning.
  • Worked in migrating HiveQL into Impala to minimize query response time.
  • Involved in loading data from edge node to HDFS using shell scripting.
  • Worked with Kerberos and integrated it to the Hadoop cluster to make it more strong and secure from unauthorized access.
  • Created Hive tables, dynamic partitions, buckets for sampling, and analysed them using HQL.
  • Created POC using Kafka and HBase for processing streaming data.
  • Involved in advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
  • Ingested huge amounts of data into HDFS using Apache Kafka.
  • Implemented automated jobs to load data from different sources and integrated with Kafka.
  • Integrated Oozie with rest of the Hadoop stack to manage Hadoop jobs (such as Map-Reduce, Pig, Hive, and SQOOP) as well as system specific jobs (such as Java programs and shell scripts).

Environment: Map Reduce, HDFS, Spark, Scala, Kafka, Hive, Pig, Spark streaming, HBase, maven, Jenkins, UNIX, Python, MRUnit, Git.

Confidential, Irving, TX

HADOOP DEVELOPER

Responsibilities:
  • Involved in getting huge volumes of data from MySQL and file servers to store in Hive.
  • Involved in initiating and successfully completing proof of concept on flume for pre-processing.
  • Experience on flume to collect log data from various sources and transfer data into Hive tables using different SerDe’s to store Json, Parquet, xml, and Avro file formats.
  • Worked with different Hive file formats like RC file, Sequence file, ORC file format and Parquet.
  • Worked on various performance optimizations like using distributed caches for small data, dynamic partitioning, Bucketing in Hive and Map Side Joins.
  • Developed Map Reduce applications using Hadoop Map-Reduce programming framework for processing.
  • Developed the code for Importing and exporting data into HDFS and Hive using SQOOP.
  • Responsible for writing Hive Queries for analyzing data in Hive warehouse using Hive Query Language (HQL).
  • Involved in defining job flows using Oozie to manage apache Hadoop jobs.
  • Wrote PIG and Hive jobs to parse the logs and structure them in tabular format to facilitate effective querying on the log data. Developed Hive User Defined Functions in Python and executed them with Hive Queries.
  • Experienced in managing and reviewing Hadoop log files.
  • Responsible to manage data coming from different sources, and perform various transformations like filtering data, joining tables and aggregations.
  • Designed, documented operational problems by following standards and procedures using JIRA.

Environment: Hive, SQL, Pig, Flume, Kafka, Map reduce, SQOOP, Scala, Python, Java, Shell Scripting, Unix Scripting, Spark, Teradata, MySQL, Oozie.

Confidential, Roseland, NJ

ETL DEVELOPER

Responsibilities:

  • Gather requirements from business users and analyse data based on the requirements.
  • Extensively used flat files and developed complex Informatica mappings using expressions, aggregators, filters, lookup and stored procedures to ensure movement of the data between various applications.
  • Designed and developed complex mappings using various transformations in Designer to extract the data from sources like Oracle, SQL Server and flat files to perform mappings based on company requirements and load into Oracle tables.
  • Worked with DQ Architect in understanding the current state of Data.
  • Data Profiling, Data Cleansing, Data Standardization, Data De-Duplication using Informatica Data Quality.
  • Generation and maintaining of IDQ workflows for defined rules.
  • Provided Knowledge Transfer to the end users.
  • Involved in functional and technical design documentation sessions with Technical team and business users.

Confidential, Bloomington, IL

ETL DEVELOPER

Responsibilities:
  • Performance tuning of SQL and PLSQL Scripts.
  • Involved in technical discussions for Business requirement, interaction with business analysts to gather the requirements.
  • Writing PL/SQL scripts, procedures and packages.
  • Performed data loads using SQL LOADER.
  • Tuned SQL and PL/SQL code using Oracle Enterprise Manager Tools and Explain plan.
  • Supported Oracle developers. Performed database tuning, created database reorganization procedures, scripted database alerts, and monitored scripts.
  • Maintaining data integrity.
  • Created partitions and aggregations.
  • Performed efficient tuning of SQL queries and stored procedures based on query execution plans, statistics and profiling.
  • Participated in regular project status meetings.

We'd love your feedback!