Data Engineer Resume
Objective
- To obtain a challenging position in the field of information technology this will utilize and refine my technologies and managerial skills.
SUMMARY
- Highly dedicated, inspiring, and expert ETL Data Engineer with overn6 years of IT industry experience exploring various technologies, tools and databases like Big Data, AWS, S3, Snowflake, Hadoop.
- Over 6+ years of overall IT experience in a variety of industries, which includes hands on experience in Big Data and Data warehouse ETL technologies.
- Experience in Extraction, Transformation and Loading (ETL) data from various sources into Data Warehouses, as well as data processing like collecting, aggregating, and moving data from various sources using.
- Good working knowledge on Clou DB, Snowflake and Teradata databases.
- Excellent Programming skills at a higher level of abstraction using Scala and Python.
- Strong experience and knowledge of real time data analytics using Spark Streaming, Kafka, and Flume.
- Working knowledge of Amazon’s Elastic Cloud Compute (EC2) infrastructure for computational tasks and Simple Storage Service (S3) as Storage mechanism.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs.
- Used IDEs like Eclipse, IntelliJ IDE, PyCharm IDE, Notepad ++, and Visual Studio for development.
- Experience working with GitHub/Git 2.12 source and version control systems.
- Improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark - SQL, Data Frame, Pair RDD's, YARN.
- Hands on experience in handling Hive tables using Spark SQL.
- Efficient in writing MapReduce Programs and using Apache Hadoop API for analyzing the structured and unstructured data.
- Developed and designed automation framework using Python and Shell scripting
- Experience in AWS EC2, configuring the servers for Auto scaling and Elastic load balancing.
- Experience with all stages of the SDLC and Agile Development model right from the requirement gathering to Deployment and production support.
- Involved in daily SCRUM meetings to discuss the development/progress and was active in making scrum meetings more productive.
TECHNICAL SKILLS
Languages: Python, Scala, C, SQL, PL/SQL, XML, Unix Shell Script
Data Modeling: Star Schema Modelling, Snowflake Modelling
Databases: Oracle11g/10g, DB2, MS SQL Server 2014/2012/2008 , Teradata
Data Warehousing / ETL Tool: Informatica 10.x,9.x (Repository Admin Console, Repository Manager, Designer, Workflow Manger, Workflow Monitor, Power Exchange), Informatica IDQ, Informatica MDM, Informatica Big Data Edition (BDE/BDM), Oracle Warehouse Builder, DataStage, SSIS, SSRS, Azure
Hadoop Eco System: HDFS, YARN, MapReduce, Apache Pig, Hive, Flume, Sqoop, Oozie, Impala and Kafka, Real - time processing Framework (Apache Spark)
Cloud Base: AWS, S3, AWS Glue, Athena, Lambda, Rest API, SOAP, GCP
Tools: Toad, SQL* plus, Putty, Tidal, Autosys, MicroStrategy, Tableau, Power BI, Cognos, GIT
Operating Systems: Windows XP 10/8/7, UNIX & Linux
PROFESSIONAL EXPERIENCE
Confidential
Data Engineer
Responsibilities:
- Worked on AWS Data pipeline to configure data loads from S3 to into Redshift.
- Using AWS Redshift Extracted, transformed, and loaded data from various heterogeneous data sources and destinations
- Created Tables, Stored Procedures, and extracted data using T-SQL for business users whenever required.
- Performs data analysis and design, and creates and maintains large, complex logical and physical data models, and metadata repositories using ERWIN and MB MDR
- I have written shell script to trigger data Stage jobs.
- Assist service developers in finding relevant content in the existing reference models.
- Like Access, Excel, CSV, Oracle, flat files using connectors, tasks and transformations provided by AWS Data Pipeline.
- Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.
- Worked on developing Pyspark script to encrypting the raw data by using hashing algorithms concepts on client specified columns.
- Responsible for Design, Development, and testing of the database and Developed Stored Procedures, Views, and Triggers
- Developed Python-based API (RESTful Web Service) to track revenue and perform revenue analysis.
- Compiling and validating data from all departments and Presenting to Director Operation.
- KPI calculator Sheet and maintain that sheet within SharePoint.
- Created Tableau reports with complex calculations and worked on Ad-hoc reporting using PowerBI.
- Creating data model that correlates all the metrics and gives a valuable output.
- Worked on the tuning of SQL Queries to bring down run time by working on Indexes and Execution Plan.
- Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
- Pre-processing using Hive and Pig.
Environment: Informatica PowerCenter 10.1, Python, Hive, Hadoop, Informatica BDM, S3, Spark, TOAD, PL/SQL, Hive, Impala, Parquet Files, Hue, Tableau BI Reports, SQL Server, Oracle 12/11g, Windows and Unix, Shell Scripting, JIRA
Confidential
ETL Data Engineer
Responsibilities:
- Worked on multiple different projects for the migration of data from on-premises to cloud
- Worked on Scala code base related to Apache Spark performing the Actions, Transformations n RDDs, Data Frames & Datasets using SparkSQL and Spark Streaming Contexts
- Designed and deployed multi-tier applications using AWS services focusing on high availability, fault tolerance, and auto-scaling in AWS Cloud Formation
- Worked on ETL pipeline to source tables and deliver calculated ratio data
- Experience in using and tuning relational databases (e.g., Microsoft SQL Server, Oracle,
- MySQL) and columnar databases (e.g., Amazon Redshift, Microsoft SQL Data Warehouse)
- Developed and documented ETL strategy to populate Data Warehouse from source systems
- Created Azure incremental pipelines or data analysis.
- Extracted data from Google Analytics using Python to RDBMS and later moved to Azure Storage
- Employed Agile methodology for project management
- Used Spark and SparkSQL to read parquet data and create tables in Hive using Scala API
- Handled importing of data from various data sources and performed transformations using Hive and MapReduce; loaded data into HDFS and Extracted data from SQL into HDFS using Sqoop.
- Wrote UNIX shell scripts for cleanup, error logging, text parsing, job scheduling and sequencing.
- Performed backend testing of procedures, functions, packages, and triggers
Environment: Informatica Power Center 9.X, AWS, S3, Redshift, IDQ, Oracle 10g/11g, Putty, SQL Server, PL-SQL, TOAD WinSCP, T-SQL, UNIX Shell Script, Tidal scheduler, DataStage, PowerBI, Autosys
ConfidentialETL Informatica Developer
Responsibilities:
- Extraction, Transformation and Loading of the data using Informatica.
- Designed the target load process based on the requirements.
- Enhancing the existing mappings where changes are made to the existing mappings using Informatica Power center.
- Develop Mappings and Workflows to load the data into Oracle tables.
- Developed various transformations like Source Qualifier, Update Strategy, Lookup transformation, Expressions and Sequence Generator for loading the data into target table.
- Created Workflows, Tasks, database connections using Workflow Manager.
- Developed Informatica mappings and tuned them for better performance.
- Created sessions and batches to move data at specific intervals & on demand using Server Manager.
- Responsibilities include creating the sessions and scheduling the sessions.
- Plan and implement migration strategy to move Oracle RDBMS from HP-UX to Linux.
- Plan, install, configure, and test high availability solution such as Oracle 10g Data Guard, and Oracle 11g and 10g RAC.
- Involved in extracting the data from SQL Server and Flat files.
- Implemented performance tuning techniques by identifying and resolving the bottlenecks in source, target, transformations, mappings, and sessions to improve performance Understanding the Functional Requirements.
- Responsible for identifying the missed records in different stages from source to target and resolving the issue.
- Extensively worked on the performance tuning for mappings and ETL procedures both at mapping and session level.
- Good experience in UNIX working environment.
Environment: Informatica Power center 9.1/8.6, SQL * Plus, TOAD, UNIX, Oracle11g/10g, MS SQL, Business Objects.
