We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

Objective

  • To obtain a challenging position in the field of information technology this will utilize and refine my technologies and managerial skills.

SUMMARY

  • Highly dedicated, inspiring, and expert ETL Data Engineer with overn6 years of IT industry experience exploring various technologies, tools and databases like Big Data, AWS, S3, Snowflake, Hadoop.
  • Over 6+ years of overall IT experience in a variety of industries, which includes hands on experience in Big Data and Data warehouse ETL technologies.
  • Experience in Extraction, Transformation and Loading (ETL) data from various sources into Data Warehouses, as well as data processing like collecting, aggregating, and moving data from various sources using.
  • Good working knowledge on Clou DB, Snowflake and Teradata databases.
  • Excellent Programming skills at a higher level of abstraction using Scala and Python.
  • Strong experience and knowledge of real time data analytics using Spark Streaming, Kafka, and Flume.
  • Working knowledge of Amazon’s Elastic Cloud Compute (EC2) infrastructure for computational tasks and Simple Storage Service (S3) as Storage mechanism.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs.
  • Used IDEs like Eclipse, IntelliJ IDE, PyCharm IDE, Notepad ++, and Visual Studio for development.
  • Experience working with GitHub/Git 2.12 source and version control systems.
  • Improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark - SQL, Data Frame, Pair RDD's, YARN.
  • Hands on experience in handling Hive tables using Spark SQL.
  • Efficient in writing MapReduce Programs and using Apache Hadoop API for analyzing the structured and unstructured data.
  • Developed and designed automation framework using Python and Shell scripting
  • Experience in AWS EC2, configuring the servers for Auto scaling and Elastic load balancing.
  • Experience with all stages of the SDLC and Agile Development model right from the requirement gathering to Deployment and production support.
  • Involved in daily SCRUM meetings to discuss the development/progress and was active in making scrum meetings more productive.

TECHNICAL SKILLS

Languages: Python, Scala, C, SQL, PL/SQL, XML, Unix Shell Script

Data Modeling: Star Schema Modelling, Snowflake Modelling

Databases: Oracle11g/10g, DB2, MS SQL Server 2014/2012/2008 , Teradata

Data Warehousing / ETL Tool: Informatica 10.x,9.x (Repository Admin Console, Repository Manager, Designer, Workflow Manger, Workflow Monitor, Power Exchange), Informatica IDQ, Informatica MDM, Informatica Big Data Edition (BDE/BDM), Oracle Warehouse Builder, DataStage, SSIS, SSRS, Azure

Hadoop Eco System: HDFS, YARN, MapReduce, Apache Pig, Hive, Flume, Sqoop, Oozie, Impala and Kafka, Real - time processing Framework (Apache Spark)

Cloud Base: AWS, S3, AWS Glue, Athena, Lambda, Rest API, SOAP, GCP

Tools: Toad, SQL* plus, Putty, Tidal, Autosys, MicroStrategy, Tableau, Power BI, Cognos, GIT

Operating Systems: Windows XP 10/8/7, UNIX & Linux

PROFESSIONAL EXPERIENCE

Confidential

Data Engineer

Responsibilities:

  • Worked on AWS Data pipeline to configure data loads from S3 to into Redshift.
  • Using AWS Redshift Extracted, transformed, and loaded data from various heterogeneous data sources and destinations
  • Created Tables, Stored Procedures, and extracted data using T-SQL for business users whenever required.
  • Performs data analysis and design, and creates and maintains large, complex logical and physical data models, and metadata repositories using ERWIN and MB MDR
  • I have written shell script to trigger data Stage jobs.
  • Assist service developers in finding relevant content in the existing reference models.
  • Like Access, Excel, CSV, Oracle, flat files using connectors, tasks and transformations provided by AWS Data Pipeline.
  • Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.
  • Worked on developing Pyspark script to encrypting the raw data by using hashing algorithms concepts on client specified columns.
  • Responsible for Design, Development, and testing of the database and Developed Stored Procedures, Views, and Triggers
  • Developed Python-based API (RESTful Web Service) to track revenue and perform revenue analysis.
  • Compiling and validating data from all departments and Presenting to Director Operation.
  • KPI calculator Sheet and maintain that sheet within SharePoint.
  • Created Tableau reports with complex calculations and worked on Ad-hoc reporting using PowerBI.
  • Creating data model that correlates all the metrics and gives a valuable output.
  • Worked on the tuning of SQL Queries to bring down run time by working on Indexes and Execution Plan.
  • Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
  • Pre-processing using Hive and Pig.

Environment: Informatica PowerCenter 10.1, Python, Hive, Hadoop, Informatica BDM, S3, Spark, TOAD, PL/SQL, Hive, Impala, Parquet Files, Hue, Tableau BI Reports, SQL Server, Oracle 12/11g, Windows and Unix, Shell Scripting, JIRA

Confidential

ETL Data Engineer

Responsibilities:

  • Worked on multiple different projects for the migration of data from on-premises to cloud
  • Worked on Scala code base related to Apache Spark performing the Actions, Transformations n RDDs, Data Frames & Datasets using SparkSQL and Spark Streaming Contexts
  • Designed and deployed multi-tier applications using AWS services focusing on high availability, fault tolerance, and auto-scaling in AWS Cloud Formation
  • Worked on ETL pipeline to source tables and deliver calculated ratio data
  • Experience in using and tuning relational databases (e.g., Microsoft SQL Server, Oracle,
  • MySQL) and columnar databases (e.g., Amazon Redshift, Microsoft SQL Data Warehouse)
  • Developed and documented ETL strategy to populate Data Warehouse from source systems
  • Created Azure incremental pipelines or data analysis.
  • Extracted data from Google Analytics using Python to RDBMS and later moved to Azure Storage
  • Employed Agile methodology for project management
  • Used Spark and SparkSQL to read parquet data and create tables in Hive using Scala API
  • Handled importing of data from various data sources and performed transformations using Hive and MapReduce; loaded data into HDFS and Extracted data from SQL into HDFS using Sqoop.
  • Wrote UNIX shell scripts for cleanup, error logging, text parsing, job scheduling and sequencing.
  • Performed backend testing of procedures, functions, packages, and triggers

Environment: Informatica Power Center 9.X, AWS, S3, Redshift, IDQ, Oracle 10g/11g, Putty, SQL Server, PL-SQL, TOAD WinSCP, T-SQL, UNIX Shell Script, Tidal scheduler, DataStage, PowerBI, Autosys

Confidential

ETL Informatica Developer

Responsibilities:

  • Extraction, Transformation and Loading of the data using Informatica.
  • Designed the target load process based on the requirements.
  • Enhancing the existing mappings where changes are made to the existing mappings using Informatica Power center.
  • Develop Mappings and Workflows to load the data into Oracle tables.
  • Developed various transformations like Source Qualifier, Update Strategy, Lookup transformation, Expressions and Sequence Generator for loading the data into target table.
  • Created Workflows, Tasks, database connections using Workflow Manager.
  • Developed Informatica mappings and tuned them for better performance.
  • Created sessions and batches to move data at specific intervals & on demand using Server Manager.
  • Responsibilities include creating the sessions and scheduling the sessions.
  • Plan and implement migration strategy to move Oracle RDBMS from HP-UX to Linux.
  • Plan, install, configure, and test high availability solution such as Oracle 10g Data Guard, and Oracle 11g and 10g RAC.
  • Involved in extracting the data from SQL Server and Flat files.
  • Implemented performance tuning techniques by identifying and resolving the bottlenecks in source, target, transformations, mappings, and sessions to improve performance Understanding the Functional Requirements.
  • Responsible for identifying the missed records in different stages from source to target and resolving the issue.
  • Extensively worked on the performance tuning for mappings and ETL procedures both at mapping and session level.
  • Good experience in UNIX working environment.

Environment: Informatica Power center 9.1/8.6, SQL * Plus, TOAD, UNIX, Oracle11g/10g, MS SQL, Business Objects.

We'd love your feedback!