Data Engineer Resume
Saint Louis, MO
SUMMARY:
- 7 years of overall IT experience in a variety of industries, which includes hands on experience of 4.5 years in Big Data technologies and designing and implementing Map Reduce.
- Excellent knowledge on Hadoop ecosystems such as HDFS, Name Node, Data Node, Resource Manager, Node Manager, Application Master and Map Reduce programming paradigm.
- Hands on experience in Capturing data from existing relational databases (MySQL, SQL SERVER and Teradata) that provide SQL interfaces using Sqoop.
- Hands on experience in working with different file formats and Combiners, Counters, Dynamic Partitions, Bucketing for best practice and performance improvement.
- Developed multiple Map - Reduce jobs to perform data cleaning and preprocessing.
- Designed HIVE queries & Pig scripts to perform data analysis, data transfer and table design to load data into Hadoop environment.
- Building and improving ETL process for loading data warehouse.
- Expertise in writing Hive UDF, Generic UDF's to in corporate complex business logic into Hive Queries.
- Used Spark API over Hortonworks Hadoop YARN to perform analytics on data.
- Exposure in working with data frames.
- Performed different ETL operations using Pig for joining operations and transformations on data to join, clean, aggregate and analyze data.
- Involved In working with Maven, Ant, sbt and Gradle for build process.
- Extensive experience on importing and exporting data using stream processing platforms like Flume and Kafka
- Experience in data workflow scheduler Zoo-Keeper and Oozie to manage Hadoop jobs by Direct Acyclic Graph (DAG) of actions with the control flows.
- Knowledge in creating impala views on top of Hive tables for faster access to analyze data.
- Integrated BI tool like Tableau with Impala and analyzed the data.
- Experienced in performance tuning and real-time analytics in both relational database and NoSQL database (HBase).
- Building and improving ETL process for loading data warehouse.
- Analyze and process complex data sets using advanced querying, visualization and analytics tools.
- Involved in Agile methodologies, daily scrum meetings, spring planning.
- Good analytical, communication, problem solving skills and adore learning new technical, functional skills.
TECHNICAL SKILLS:
Bigdata Ecosystem: HDFS and Map Reduce, Pig, Hive, Impala, YARN, HUE, Oozie, Zookeeper, Solr, Apache Spark, Apache STORM, Apache Kafka, Sqoop, Flume.
NoSQL Databases: HBase, Cassandra, and MongoDB
Hadoop Distributions: Cloudera, Hortonworks
Programming languages: Java, C/C++, SCALA, Pig Latin, HiveQL.
Scripting Languages: Shell Scripting, Java Scripting
Databases: MySQL, oracle, Teradata, DB2
Build Tools: Maven, Ant, Gradle, sbt
Reporting Tool: Tableau
Version control Tools: SVN, Git, GitHub
Cloud: AWS, Azure
App/Web servers: WebSphere, WebLogic, JBoss and Tomcat
Web Design Tools: HTML, AJAX, JavaScript, JQuery, CSS and JSON.
Operating Systems: WINDOWS 10/8/Vista/ XP
Development IDEs: NetBeans, Eclipse IDE, Python(IDLE)
Packages: Microsoft Office, putty, MS Visual Studio
WORK EXPERIENCE:
Data Engineer
Confidential, Saint Louis, MO
Technologies/tools: Apache Hadoop, HDFS, MapReduce, Sqoop, Flume, Pig, Hive, HBase, Oozie, Scala, Spark
Responsibilities:
- Extracted the data from RDBMS into HDFS using Sqoop.
- Used Kafka to collect, aggregate and store the web log data from different sources like web servers, mobile and network devices and pushed into HDFS.
- Implemented MapReduce programs on log data to transform into structured way to find user information.
- Used Pig scripts to transform raw data from several data sources into forming baseline data.
- Developed UDF functions for Hive and wrote complex queries in Hive for data analysis.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
- Used ETL processes to load data from flat files into the target database by applying business logic on transformation mapping for inserting and updating records when loaded.
- Worked on data serialization formats for converting complex objects into sequence bits by using Avro, RC and ORC file formats.
- Created and worked Sqoop jobs with incremental load to populate Hive External tables.
- Designed and developed Hive tables to store staging and historical data.
- Created Hive tables as per requirement, internal and external tables are defined with appropriate static and dynamic partitions, intended for efficiency.
- Processed large data sets utilizing Hadoop cluster. The data that are stored on HDFS were preprocessed/validated using Pig and then processed data was stored into Hive warehouse which enabled Business analysts to get the required data from Hive.
- Experience in using ORC file format with Snappy compression for optimized storage of Hive tables.
- Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation.
- Developed Oozie workflow for scheduling and orchestrating the ETL process.
- Involved in migrating MapReduce jobs into Spark jobs and used Spark SQL and Data Frames API to load structured and semi-structured data into Spark clusters.
- Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing.
- Implemented Spark using Scala and Spark SQL for faster testing and processing of data.
- Used Apache Kafka for importing real time network log data into HDFS.
- Integrated BI tool like Tableau with Impala and analyzed the data.
- Converted event streams into structured data using Confidential Cloud Dataflow (Apache Beam and Java).
Senior Software Engineer
Confidential, Vernon Hills, IL
Technologies/tools: Hadoop, Map Reduce, Hive, Java, Maven, Impala, Pig, Spark, Oozie, Oracle, Yarn, GitHub, Tableau, Unix, Cloudera, Sqoop, Scala, HBase.
Responsibilities:
- Developed SQOOP scripts for importing and exporting data into HDFS and Hive.
- Developing design documents considering all possible approaches and identifying best of them.
- Developed the services to run the Map-Reduce jobs as per the requirement basis.
- Responsible to manage data coming from different sources.
- Developing business logic using Scala.
- Responsible for loading data from UNIX file systems to HDFS. Installed and configured Hive and written Pig/Hive UDFs.
- Involved in creating Hive Tables, loading with data and writing Hive queries which will invoke and run MapReduce jobs in the backend.
- Writing MapReduce programs to convert text files into AVRO and loading into Hive tables.
- Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi structured data coming from various sources.
- Developed scripts and automated data management from end to end and sync up between all the clusters.
- Exploring with Spark for improving the performance and optimization of the existing algorithms in Hadoop.
- Import the data from different sources like HDFS/HBase into Spark RDD.
- Experienced with Spark Context, Spark -SQL, Data Frame, Pair RDD's, Spark YARN.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDD and Scala.
- Implemented the workflows using Apache Oozie framework to automate tasks.
- Imported results into visualization BI tool Tableau to create dashboards.
- Worked in Agile Methodology and used JIRA for maintain the stories about project.
- Involved in gathering the requirements, designing, development and testing.
Software Engineer
Confidential
Technologies/tools: includes DataStage, Teradata, Control M, Cognos, Framework Manager
Responsibilities:
- Worked as Software Engineer at Confidential Global Technologies and was involved in HSBC’s Global Information System (GIS).
- Delivered global BI solutions for various functional units of bank.
- Experienced in Cognos10 with Report Studio, Framework Manager, Query Studio, Analysis Studio.
- Experience in Designing, Developing, Documenting, Testing of ETL jobs and mappings in Server and Parallel jobs using DataStage to populate tables in Data Warehouse and Data marts.
- Expertise in creating databases, users, tables, triggers, macros, views, stored procedures, functions, Packages, joins and hash indexes in Teradata database.
- Design, develop, implement and maintain new reporting functionality and analytic applications across multiple business units using various business intelligence tools
- Implementation of data quality improvement processes and initiatives.
- Strong understanding of dimensional modeling of Star Schemas, Snow-Flake Schemas Methodologies for building enterprise data warehouses.
Quality Analyst
Confidential
Responsibilities:
- Worked as Quality Analyst at Confidential and was involved in Landing Page Quality (LPQ).
- Worked on internal tools to collect data from different sources and maintain Quality of data by applying different business rules.
