We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

MI

SUMMARY

  • 8 years of technical experience in Data engineer/Modelling business needs of clients, developing effective and efficient solutions, and ensuring client deliverables within committed timelines.
  • Deep knowledge and strong deployment experience in Hadoop and Big Data ecosystems - HDFS, MapReduce, Spark, Sqoop, Hive, Kafka, zookeeper, and HBase.
  • Extensively worked on Spark using Scala on cluster for computational analytics, installed it on top of Hadoop performed advanced analytical application by making use of Spark with Hive and SQL/Oracle.
  • Experienced Data Modeler with strong conceptual, Logical and Physical Data Modelling skills, Data Profiling skills, Maintaining Data Quality, experience with JAD sessions for requirements gathering, creating data mapping documents, writing functional specifications, queries. And Dimensional Data Modelling, FACT & Dimension tables.
  • Expertise in AWS Resources like EC2, S3, EBS, VPC, ELB, SNS, RDS, IAM, Route 53, Auto scaling, Cloud Formation, Cloud Watch, Security Groups.
  • Skilful in Data Analysis using SQL on Oracle, MS SQL Server, DB2, Teradata and AWS.
  • Experienced in trouble shooting ETL jobs, Datawarehouse, Data mart data store models.
  • Assisted in creating communication materials based on data for key internal / external audiences.
  • Expert in documenting the Business Requirements Document (BRD), generating the UAT Test Plans, maintaining the Traceability Matrix and assisting in Post Implementation activities.
  • Enterprise Data Modeler with a deep understanding of developing Enterprise Data Models that strictly meet Normalization Rules, as well as Enterprise Data Warehouses using Kimball and Data Warehouse Methodologies.
  • Knowledgeable in Best Practices and Design Patterns, Cube design, BI Strategy and Design and 3NFModeling.
  • Delivered zero defect code for three large projects which involved changes to both front end (web services) and back-end (Oracle, Teradata).
  • Constructing and manipulating large datasets of structured, semi-structured, and unstructured data and supporting systems application architecture using tools like SAS, SQL, Python, R, Minitab, PowerBI, and more to extract multi-factor interactions and drive change.
  • Using the S3 CLI tools, create scripts for creating new snapshots and deleting existing snapshots in S3.
  • Expertise with Hive data warehouse architecture, including table creation, data distribution using Partitioning and Bucketing, and query development and tuning in Hive Query Language.
  • Experienced in leading the enhancement, architecture, and ongoing evolution using a wide array of technologies (Spark, Python, Apigee, Delta Lake, Databricks, Kafka, Data bucks, as well as more traditional technologies such as MuleSoft and SQL) across Amazon cloud environment.
  • Knowledge of job workflow management and monitoring tools like Oozie, and zookeeper.
  • Proficient in writing Bash, Pearl, Python scripts to automate and provide Control Flow.
  • I have handled performance tuning, conducted backups, and ensure integrity and security of databases managed Postgres DB in the AWS environment and Aurora-Postgres.
  • Familiar with Data Stage production job scheduling.

PROFESSIONAL EXPERIENCE

Confidential, MI

Data Engineer

Responsibilities:

  • Involved in Data mapping specifications to create and execute detailed system test plans. The data mapping specifies what data will be extracted from an internal data warehouse, transformed, and sent to an external entity.
  • Documented logical, physical, relational, and dimensional data models. Designed the Data Marts in dimensional data modelling using star and snowflake schemas.
  • Prepared documentation for all entities, attributes, data relationships, primary and foreign key structures, allowed values, codes, business rules, and glossary evolve and change during the project.
  • Coordinated with DBA on database build and table normalizations and de-normalizations.
  • Identified the entities and relationship between the entities to develop Conceptual Model busing ERWIN.
  • Involved with Data Profiling activities for new sources before creating new subject areas in warehouse.
  • Extensively worked Data Governance, i.e., Metadata management, Master data Management, Data Quality, Data Security.
  • Enforced referential integrity in the OLTP data model for consistent relationship between tables and efficient database design.
  • Tested the ETL process for both before data validation and after data validation process. Tested the messages published by ETL tool and data loaded into various databases.
  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of AWS, T-SQL, Spark SQL, and U-SQL.
  • Designed and implemented scalable infrastructure and platform for large amounts of data ingestion, aggregation, integration, and analytics in Hadoop, including Spark, Hive, HBase.
  • Creating MapReduce programs to enable data for transformation, extraction, and aggregation of multiple formats like Avro, Parquet, XML, JSON, CSV, and other compressed file formats.
  • Developing an architecture to move the project from Abinitio to py spark and Scala spark.
  • Use Python, Scala programming daily to perform transformations for applying business logic.
  • Worked with Informatica Cloud for data integration between Salesforce, RightNow, Eloqua, and Web Services applications.
  • Created Pipelines in ADF using Linked Services, Datasets, Pipeline to Extract, Transform, and load data from different sources like SQL, SQL Data warehouse, write-back tool, and backward.
  • Using Enterprise data lake to support various use cases including Analytics, Storing, and reporting of Voluminous, structured, and unstructured, rapidly changing data.
  • Using Sqoop to load data from HDFS, Hive, MySQL, and many other sources on daily basis.
  • Exported the analysed data into relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Converting data load pipeline algorithms written in python and SQL to Scala spark and pyspark.

Confidential, Greenwood Village, CO

Data Engineer

Responsibilities:

  • Involved in Data mapping specifications to create and execute detailed system test plans. The data mapping specifies what data will be extracted from an internal data warehouse, transformed, and sent to an external entity.
  • Documented logical, physical, relational, and dimensional data models. Designed the Data Marts in dimensional data modelling using star and snowflake schemas.
  • Involved in building scalable distributed data lake system for Confidential real time and batch analytical needs.
  • Involved in designing, reviewing, optimizing data transformation processes using Apache Storm.
  • Used Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Developed Scala scripts, UDFs using both Data frames/SQL and RDD/MapReduce in Spark 1.6 for Data Aggregation, queries and writing data back into OLTP system through Scoop.
  • Optimizing of existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames and Pair RDD’s.
  • Imported data from Kafka Consumer into HBase using Spark streaming.
  • Used Oozie workflow engine to manage interdependent Hadoop jobs and to automate several types of Hadoop jobs such as Java MapReduce, Hive and Sqoop as well as system specific jobs.
  • Experienced in handling large datasets using partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective efficient Joins, Transformation and other during ingestion process itself.
  • Prepared documentation for all entities, attributes, data relationships, primary and foreign key structures, allowed values, codes, business rules, and glossary evolve and change during the project.
  • Coordinated with DBA on database build and table normalizations and de-normalizations.
  • Identified the entities and relationship between the entities to develop Conceptual Model busing ERWIN.
  • Developed Logical Model from the conceptual model.
  • Responsible for different Data mapping activities from Source systems.
  • Involved with Data Profiling activities for new sources before creating new subject areas in warehouse.
  • Extensively worked Data Governance, i.e., Metadata management, Master data Management, Data Quality, Data Security.
  • Performed complex data analysis in support of ad-hoc and standing customer requests.
  • Enforced referential integrity in the OLTP data model for consistent relationship between tables and efficient database design.
  • Experience in creating UNIX scripts for file transfer and file manipulation.

Confidential, Lake Success, NY

Data Analyst

Responsibilities:

  • Analysed business requirements, system requirements, data mapping requirement specifications, and responsible for documenting functional requirements and supplementary requirements in Quality Centre.
  • Setting up of environments to be used for testing and the range of functionalities to be tested as per technical specifications.
  • Tested Complex ETL Mapping and Sessions based on business user requirements and business rules to load data from source flat files and RDBMS tables to target tables.
  • Responsible for different Data mapping activities from Source systems to EDW, ODS& data marts.
  • Delivered file in various file formatting system (ex. Excel file, Tab delimited text, Coma separated text, Pipe delimited text etc.)
  • Performed ad hoc analyses, as needed, with the ability to comprehend analysis as needed.
  • Involved in Teradata SQL Development, Unit testing and Performance tuning and to ensure testing issues are resolved based on using defect reports.
  • Tested the database to check field size validation, check constraints, stored procedures and cross verifying the field size defined within the application with metadata.
  • Installed, designed, and developed the SQL Server database.
  • Created a logical design of the central relational database using Erwin.
  • Configured the DTS packages to run in periodic intervals.
  • Extensively worked with DTS to load the data from source systems and run-in periodic intervals.
  • Worked with data transformations in both normalized and de-normalized data environments.
  • Involved in data manipulation using stored procedures and Integration Services.
  • Worked on query optimization, stored procedures, views, and triggers.
  • Assisted in OLAP and Data Warehouse environment when assigned.
  • Created tables, views, triggers, stored procedures, and indexes.
  • Designed and implemented database replication strategies for both internal and Disaster Recovery.
  • Created ftp connections, database connections for the sources and targets.
  • Maintained security and data integrity of the database.
  • Developed several forms & reports using Crystal Reports.

Confidential, Dallas, TX

Data Engineer

Responsibilities:

  • Worked with Data Vault Methodology Developed Normalized Logical and Physical database models.
  • Developed Data Mapping, Data Governance, Transformation, and cleansing rules for the Master Data Management Architecture involving OLTP, ODS, and OLAP.
  • Worked on Performance Tuning of the database which includes indexes, optimizing SQL Statements.
  • Created tables, views, sequences, triggers, table spaces, constraints, and generated DDL scripts for physical implementation.
  • Developed mapping spreadsheets for (ETL) team with source to target data mapping with physical naming standards, datatypes, volumetric, domain definitions, and corporate meta-data definitions.
  • Established and maintained comprehensive data model documentation including detailed descriptions of business entities, attributes, and data relationships.
  • Implemented Data Vault Modelling Concept solved the problem of dealing with change in the environment by separating the business keys and the associations between those business keys, from the descriptive attributes of those keys using HUB, LINKS tables and Satellites.
  • Maintain and work with our data pipeline that transfers and processes several terabytes of data using Spark, Scala, Python, Apache Kafka, Pig/Hive & Impala
  • Apply data analysis, data mining and data engineering to present data clearly. detailed production level using Workflow Diagrams, Sequence Diagrams, Activity Diagrams and Use Case Modelling.
  • Have been working with AWS cloud services (VPC, EC2, S3, RDS, Redshift, Data Pipeline, EMR, DynamoDB, Workspaces, Lambda, Kinesis, RDS, SNS, SQS).
  • Involved in creating Physical and Logical models using Erwin.
  • Involved with Data Analysis Primarily Identifying Data Sets, Source Data, Source Meta Data, Data Definitions and Data Formats.
  • Expert in the Data Analysis, Design, Development, Implementation and Testing using Data Conversions, Extraction, Transformation and Loading (ETL) and ORACLE, SQL Server, and other relational and non-relational databases.
  • Document all data mapping and transformation processes in the Functional Design documents based on the business requirements.
  • Generated ad-hoc SQL queries using joins, database connections and transformation rules to fetch data from legacy DB2 and SQL Server database systems.
  • Highly proficient in Data Modelling retaining concepts of RDBMS, Logical and Physical Data Modelling until 3NormalForm (3NF) and Multidimensional Data Modelling Schema (Star schema, Snow-Flake Modelling, Facts, and dimensions).
  • Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, experienced in Maintaining the Hadoop on AWS EMR.
  • Generated ad-hoc SQL queries using joins, database connections and transformation rules to profile data from DB2 and SQL Server database systems.
  • Worked with data compliance teams, Data governance team to maintain data models,

Confidential

ETL Developer

Responsibilities:

  • Involved in Migrating historical as built data from Link Tracker Oracle database to TD using Abinitio.
  • Implemented historical purge process for Clickstream, order broker &link tracker to TD using Abinitio
  • Implemented the centralized graphs concept.
  • Extensively used Abinitio components like Reformate, rollup, lookup, joiner, re-defined and also developed many sub graphs
  • Abinitio Sandbox creation at both GDE level and air command level, scheduling the interdependent jobs (Abinitio deployed graphs) through UNIX wrapper template.
  • Performing tuning the Abinitio graphs
  • Sandbox creation and adding the parameters based on the requirement
  • Involved in loading the transformed data file into TD staging tables through TD Load utilities, Fast load and Multi load scripts, and Creating TD macro’s for loading the data from staging to target tables
  • Performed the data validation on TD warehouse data as per few standard test cases
  • Leading the module to load all PARTY RELSHIP tables, responsible for requirement gathering, creating specification documents & Test cases documents, designing and validating the ETL mapping, development through Unit Testing, validating the data populated in the data base and giving UAT support and Resolution of issues raised by the users and different groups.
  • Responsible as E-R consultant, ER(Extract-Replicate) Golden gate tool which is used to extract the real time data to warehouse without hitting to the database which pulls the data from oracle Archive logs as oracle 10g support as ASM (Automatic storage mechanism) method.
  • Also involved to designing the Data Allegro post scripts to load the data from LRF files to DA database.

We'd love your feedback!