We provide IT Staff Augmentation Services!

Data Engineer Resume

3.00/5 (Submit Your Rating)

Saint Louis, MO

SUMMARY:

  • 7 years of overall IT experience in a variety of industries, which includes hands on experience of 4.5 years in Big Data technologies and designing and implementing Map Reduce.
  • Excellent knowledge on Hadoop ecosystems such as HDFS, Name Node, Data Node, Resource Manager, Node Manager, Application Master and Map Reduce programming paradigm.
  • Hands on experience in Capturing data from existing relational databases (MySQL, SQL SERVER and Teradata) that provide SQL interfaces using Sqoop.
  • Hands on experience in working with different file formats and Combiners, Counters, Dynamic Partitions, Bucketing for best practice and performance improvement.
  • Developed multiple Map - Reduce jobs to perform data cleaning and preprocessing.
  • Designed HIVE queries & Pig scripts to perform data analysis, data transfer and table design to load data into Hadoop environment.
  • Building and improving ETL process for loading data warehouse.
  • Expertise in writing Hive UDF, Generic UDF's to in corporate complex business logic into Hive Queries.
  • Used Spark API over Hortonworks Hadoop YARN to perform analytics on data.
  • Exposure in working with data frames.
  • Performed different ETL operations using Pig for joining operations and transformations on data to join, clean, aggregate and analyze data.
  • Involved In working with Maven, Ant, sbt and Gradle for build process.
  • Extensive experience on importing and exporting data using stream processing platforms like Flume and Kafka
  • Experience in data workflow scheduler Zoo-Keeper and Oozie to manage Hadoop jobs by Direct Acyclic Graph (DAG) of actions with the control flows.
  • Knowledge in creating impala views on top of Hive tables for faster access to analyze data.
  • Integrated BI tool like Tableau with Impala and analyzed the data.
  • Experienced in performance tuning and real-time analytics in both relational database and NoSQL database (HBase).
  • Building and improving ETL process for loading data warehouse.
  • Analyze and process complex data sets using advanced querying, visualization and analytics tools.
  • Involved in Agile methodologies, daily scrum meetings, spring planning.
  • Good analytical, communication, problem solving skills and adore learning new technical, functional skills.

TECHNICAL SKILLS:

Bigdata Ecosystem: HDFS and Map Reduce, Pig, Hive, Impala, YARN, HUE, Oozie, Zookeeper, Solr, Apache Spark, Apache STORM, Apache Kafka, Sqoop, Flume.

NoSQL Databases: HBase, Cassandra, and MongoDB

Hadoop Distributions: Cloudera, Hortonworks

Programming languages: Java, C/C++, SCALA, Pig Latin, HiveQL.

Scripting Languages: Shell Scripting, Java Scripting

Databases: MySQL, oracle, Teradata, DB2

Build Tools: Maven, Ant, Gradle, sbt

Reporting Tool: Tableau

Version control Tools: SVN, Git, GitHub

Cloud: AWS, Azure

App/Web servers: WebSphere, WebLogic, JBoss and Tomcat

Web Design Tools: HTML, AJAX, JavaScript, JQuery, CSS and JSON.

Operating Systems: WINDOWS 10/8/Vista/ XP

Development IDEs: NetBeans, Eclipse IDE, Python(IDLE)

Packages: Microsoft Office, putty, MS Visual Studio

WORK EXPERIENCE:

Data Engineer

Confidential, Saint Louis, MO

Technologies/tools: Apache Hadoop, HDFS, MapReduce, Sqoop, Flume, Pig, Hive, HBase, Oozie, Scala, Spark

Responsibilities:

  • Extracted the data from RDBMS into HDFS using Sqoop.
  • Used Kafka to collect, aggregate and store the web log data from different sources like web servers, mobile and network devices and pushed into HDFS.
  • Implemented MapReduce programs on log data to transform into structured way to find user information.
  • Used Pig scripts to transform raw data from several data sources into forming baseline data.
  • Developed UDF functions for Hive and wrote complex queries in Hive for data analysis.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
  • Used ETL processes to load data from flat files into the target database by applying business logic on transformation mapping for inserting and updating records when loaded.
  • Worked on data serialization formats for converting complex objects into sequence bits by using Avro, RC and ORC file formats.
  • Created and worked Sqoop jobs with incremental load to populate Hive External tables.
  • Designed and developed Hive tables to store staging and historical data.
  • Created Hive tables as per requirement, internal and external tables are defined with appropriate static and dynamic partitions, intended for efficiency.
  • Processed large data sets utilizing Hadoop cluster. The data that are stored on HDFS were preprocessed/validated using Pig and then processed data was stored into Hive warehouse which enabled Business analysts to get the required data from Hive.
  • Experience in using ORC file format with Snappy compression for optimized storage of Hive tables.
  • Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation.
  • Developed Oozie workflow for scheduling and orchestrating the ETL process.
  • Involved in migrating MapReduce jobs into Spark jobs and used Spark SQL and Data Frames API to load structured and semi-structured data into Spark clusters.
  • Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing.
  • Implemented Spark using Scala and Spark SQL for faster testing and processing of data.
  • Used Apache Kafka for importing real time network log data into HDFS.
  • Integrated BI tool like Tableau with Impala and analyzed the data.
  • Converted event streams into structured data using Confidential Cloud Dataflow (Apache Beam and Java).

Senior Software Engineer

Confidential, Vernon Hills, IL

Technologies/tools: Hadoop, Map Reduce, Hive, Java, Maven, Impala, Pig, Spark, Oozie, Oracle, Yarn, GitHub, Tableau, Unix, Cloudera, Sqoop, Scala, HBase.

Responsibilities:

  • Developed SQOOP scripts for importing and exporting data into HDFS and Hive.
  • Developing design documents considering all possible approaches and identifying best of them.
  • Developed the services to run the Map-Reduce jobs as per the requirement basis.
  • Responsible to manage data coming from different sources.
  • Developing business logic using Scala.
  • Responsible for loading data from UNIX file systems to HDFS. Installed and configured Hive and written Pig/Hive UDFs.
  • Involved in creating Hive Tables, loading with data and writing Hive queries which will invoke and run MapReduce jobs in the backend.
  • Writing MapReduce programs to convert text files into AVRO and loading into Hive tables.
  • Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi structured data coming from various sources.
  • Developed scripts and automated data management from end to end and sync up between all the clusters.
  • Exploring with Spark for improving the performance and optimization of the existing algorithms in Hadoop.
  • Import the data from different sources like HDFS/HBase into Spark RDD.
  • Experienced with Spark Context, Spark -SQL, Data Frame, Pair RDD's, Spark YARN.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDD and Scala.
  • Implemented the workflows using Apache Oozie framework to automate tasks.
  • Imported results into visualization BI tool Tableau to create dashboards.
  • Worked in Agile Methodology and used JIRA for maintain the stories about project.
  • Involved in gathering the requirements, designing, development and testing.

Software Engineer

Confidential

Technologies/tools: includes DataStage, Teradata, Control M, Cognos, Framework Manager

Responsibilities:

  • Worked as Software Engineer at Confidential Global Technologies and was involved in HSBC’s Global Information System (GIS).
  • Delivered global BI solutions for various functional units of bank.
  • Experienced in Cognos10 with Report Studio, Framework Manager, Query Studio, Analysis Studio.
  • Experience in Designing, Developing, Documenting, Testing of ETL jobs and mappings in Server and Parallel jobs using DataStage to populate tables in Data Warehouse and Data marts.
  • Expertise in creating databases, users, tables, triggers, macros, views, stored procedures, functions, Packages, joins and hash indexes in Teradata database.
  • Design, develop, implement and maintain new reporting functionality and analytic applications across multiple business units using various business intelligence tools
  • Implementation of data quality improvement processes and initiatives.
  • Strong understanding of dimensional modeling of Star Schemas, Snow-Flake Schemas Methodologies for building enterprise data warehouses.

Quality Analyst

Confidential

Responsibilities:

  • Worked as Quality Analyst at Confidential and was involved in Landing Page Quality (LPQ).
  • Worked on internal tools to collect data from different sources and maintain Quality of data by applying different business rules.

We'd love your feedback!