Lead Data Engineer Resume
Dallas, TX
SUMMARY:
- About 6 years of technical and functional experience in Data warehouse implementations, Data Visualization, ETL methodology using IBM DataStage, Teradata, Oracle and DB2 in Retail, Auto and Health Insurance Domains.
- Hands on experience in writing Python Scripts for Data Extract and Data Transfer from various data sources.
- Hands on experience in installing, configuring and using Apache Hadoop ecosystem utilities such as HIVE, SQOOP, SPARK, OOZIE, FLUME and KAFKA.
- Experience developing data models and processing data through big data frameworks like Hive, HDFS and Spark to access streaming data and implement data pipelines to process real - time suggestion and recommendations.
- Strong experience in processing large amounts of structured and unstructured data, including integrating data from multiple sources.
- Strong experience in Data Analysis, Data Migration, Data Profiling, Data Cleansing, Transformation, Integration, Data Import, and Data Export using multiple ETL tools such as IBM DataStage, Informatica power center.
- Strong experience in various Microsoft Azure services like Azure App services, Resource management, API management, Azure AD, Application insights, scheduling, caching, Azure SQL, NoSQL and Auto scaling.
- Experience in Configuring and implementing Azure Data Factory to pipeline data loading from on- premise flat files data lake. Also, migrated on premise apps and database to Azure.
- Experience in Automating, Configuring and Deploying Instances on Azure environments and in Data centers.
- Experience in performance tuning by identifying bottlenecks at source, target, mapping, session or transformation.
PROFESSIONAL EXPERIENCE:
Confidential, Dallas, TX
Lead Data Engineer
Responsibilities:
- Designed and developed data pipelines using Big data tools including Airflow, Spark, Sqoop, Oozie, Hive and Kafka
- Used Sqoop to extract and load incremental and non-incremental data from RDBMS systems like oracle, teradata to HDFS.
- Created multiple pipelines in azure data factory and Airflow to move data across multiple platforms.
- Developed python code for different tasks, dependencies, SLA watcher and time sensor for jobs and automated the jobs using Airflow tool.
- Developed complex code to generate KPI’s for the items in club which can be leveraged in the key decision making.
- Built complex jobs in Azure data factory and automated them to run on timely basis to reduce cluster workload.
- Worked with cross functional teams to identify the source of truth for the data and created data profile to validate the accuracy and security of the consumed data.
- Provided indefinite support to team members in building pipelines, developing Pyspark /Scala code in Databricks, creating Power BI reports, creating database stored procedures/triggers/views etc.
- Developed Pig scripts to parse raw data, populate staging tables and store refined data in partitioned DB2.
- Designed and developed Real Time Stream Processing Application using Spark, Kafka, Scala and Hive to perform Streaming ETL and apply Machine Learning.
- Created HBase tables to load large sets of semi structured data coming from various sources.
- Created Hive queries to compare the raw data with EDW tables and performing aggregates.
- Used a combination of Python and Spark and Spark SQL context to transform Semi-Structured data and to create aggregate level tables.
- Involved in data ingestion into HDFS using Sqoop from various sources using connectors and import parameters.
- Developed data pipeline using flume, Sqoop and pig to extract the data from weblogs and store in HDFS.
Environment: DB2, Python, Microsoft Azure, Azure AD, Azure Data factory, Azure Data Factory, Azure Blob storage, Azure SQL Db, Hadoop, Hive, Spark, Pig, Sqoop, kafka, Pyspark, GIT, flat files, CSV File, MS Power BI, Tableau.
Confidential, Lansing, MI
Senior Data Engineer
Responsibilities:
- Worked with the business SMEs in an agile development approach to understand the business rules required for the analytics, with a constant feedback loop throughout the development process to refine the rules.
- Compiled data from various sources public and private databases to perform complex analysis and data manipulation for actionable results.
- Created data lake factory using pipeline and copy data (including scheduler) to load data from on-premises databases to Azure databases.
- Created application insights for DataStage jobs to provide metrics, alerts and monitoring of jobs frequently.
- Processed data from various sources using Teradata utilities (TPT, BTEQ, MLOAD, FAST LOAD and FAST EXPORT). Further data is used for Tableau reporting and Ad-hoc reporting etc.
- Developed complex Teradata SQL queries using advanced analytical features and functions of Teradata 15.
- Worked on development of data warehouse, Data Lake and ETL systems using relational and non-relational tools like SQL, No SQL.
- Built complex DataStage jobs for data transformations which were driven by complicated business rules.
- Planned and implemented the Data Marts move from DB2 UDB to Azure cloud, including designing the new data structures on Azure and re-designing the ETL jobs to optimize performance with Azure.
- Wrote PL SQL scripts and used couple of metadata tables to do Data Profiling, Referential Integrity and Data Quality before loading data into staging area.
- Worked on ETL process and handled importing data from various data sources, performed transformations.
- Extensively used store procedures and UDF functions such as scalar and inline functions for complex queries, Triggers, indexes involving multiple tables from SQL server and Hadoop data sources.
- Created databases, users, tables, triggers, macros, views, stored procedures, functions, Packages, joins and hash indexes in Teradata database. Participated Meetings, gathered requirements and supporting Analysts.
- Involved in modifying various existing packages, Procedures, functions, triggers according to new business needs.
Environment: DB2, IBM DataStage 11.3, Python, Teradata 12, Teradata SQL Assistant, Microsoft Azure, Azure AD, Data lake, Data Factory, Hadoop, Hive, Spark, JS, CVS, GIT, flat files, CSV File, MS Power BI, Tableau.
Confidential
Data warehouse Developer
Responsibilities:
- Involved in requirements gathering and source data analysis and identified business rules for Data Transformations.
- Translating Business requirements into Data Mart design coordinating with team members Creating Fact, Dimensional and Aggregate Tables and Loading Data Warehouse tables.
- Extensively used the designer to develop various parallel jobs to extract, transform, integrate and load the data into Corporate Data warehouse.
- Designed and developed the ETL jobs using Parallel Extender that distributed the incoming data concurrently across all the processors, to achieve the best performance.
- Implemented the Surrogate Key by using Key Management functionality for new rows in Data Warehouse.
- Used IBM DataStage to extract data from various sources, load data into the databases like IBM DB2, Teradata.
- Designed and developed DataStage jobs for Loading Staging Data from various sources like Oracle, DB2 DB into Data Warehouse applying business rules which consists data loads, data cleansing, and data massaging.
- Extensively used the Sequential File stage, Complex Flat File Stage, Hashed File Stage, Modify, Dataset, Filter, Funnel, Join, Lookup, Copy, Aggregator, and Change Capture during ETL development.
- Extensively worked with DataStage Parallel Extender for Parallel Processing to improve job performance while working with bulk data sources.
- Created Shared Containers to simplify ETL design and used it as a common component throughout the project.
Environment: Informatica, IBM DataStage 11.3, Python, SQL Server, Teradata 12, Teradata SQL Assistant, Microsoft Azure, Azure AD, Data lake, Data Factory, JS, CVS, GIT, flat files, CSV File, MS Power BI, Tableau.
Confidential
ETL Developer
Responsibilities:
- Prioritized Business requirements and translated those requirements into functional requirements.
- Worked with Business analysts and the DBA for requirements gathering, business analysis, designing and translated the business requirements into technical specifications to build the Enterprise data warehouse.
- Analyzed business requirements, system requirements, data mapping requirement specifications, and responsible for documenting functional requirements and supplementary requirements.
- Developed ETL procedures to ensure conformity, compliance with standards and lack of redundancy, translated business rules and functionality requirements into ETL procedures using DataStage.
- Extensively used DataStage Designer to develop various jobs to extract, cleanse, transform, integrate and load data into data warehouse.
- Designed DataStage Parallel jobs involving complex business logic, update strategies, transformations, filters, lookups and necessary source-to-target data mappings to load the target.
- Extensively used the Sequential File stage, Complex Flat File Stage, Hashed File Stage, Modify, Dataset, Filter, Funnel, Join, Lookup, Copy, Aggregator, and Change Capture during ETL development.
- Used DataStage Director to Run and Monitor the Jobs for Performance Statistics.
- Created local and shared containers based on the requirement as a reusable component and simplify the design.
- Extensively worked with database objects including tables, views, indexes, schemas, PL/SQL packages, stored procedures, functions, and triggers.
- Designed Power BI data visualization utilizing cross tabs, maps, scatter plots, pie, bar and density charts.
- Responsible for the Performance tuning of the SQL Commands with the use of the SQL Profiler and Database Engine Tuning Wizard.
- Design and develop PL/SQL packages, stored procedure, tables, views, indexes and functions.
Environment: SQL SERVER 2008, T-SQL, SSIS, SSRS, Informatica PowerCenter 9.3, Teradata SQL Assistant, ORACLE 10, PL/SQL, DB2, SYBASE, JAVA, XSLT, XPATH, CSS, SOAP, WSDL, Web Services, WINDOWS, Ant 1.9.6, Maven.
Confidential
Data Analyst intern
Responsibilities:
- Web-based implementation of a recommendation system suited to business of travel agencies and banks.
- Designed and developed apps using JSP, HTML5, CSS3, JavaScript (validations), Bootstrap and AngularJS on frontend. Also developed Custom Tags, JSTL to support custom User Interfaces.
- Involved in extensive DATA validation by writing several complex SQL queries and Involved in back-end testing and worked with data quality issues.
- Creating complex Transact SQL (TSQL) queries, Sub queries, Correlated subqueries, Dynamic SQL queries to filter bad data and have a data consistency.
- Responsible for Extraction, Transformation and Loading (ETL) of data from multiple upstream sources to Database. Upstream sources are OLTP, excel, SharePoint lists & Azure SQL Database.
- Developed Power Pivots when aggregations are needed on top of few million records. Mostly these were used as Proof of concept reports by the Client.
- Created SSIS packages to extract data from OLTP to OLAP systems and Scheduled Jobs to call the packages and Stored Procedures.
- Developed Excel reports (Regular, PowerPivot). Gave a demo on using Power View to the client.
Implemented cascading parameters in SSRS Reports.
Environment: MYSQL, DB2, SSIS, SSRS, Informatica Power Center 8.x, DataStage, Oracle 11g, Tableau 9.3, SQL Developer
