We provide IT Staff Augmentation Services!

Senior Data Engineer Resume

0/5 (Submit Your Rating)

SUMMARY

  • Around 9 years of IT experience in Data Engineering and Data Modeling with high proficiency in using Big Data technologies and cloud infrastructure.
  • Good understanding and knowledge with Agile and Waterfall environments.
  • Excellent knowledge in Migrating servers, databases, and applications from on premise to Azure and Google Cloud Platform.
  • Extensive knowledge of Big Data, Hadoop, MapReduce, Hive and other emerging technologies.
  • Developing data pipeline using Sqoop, and MapReduce to ingest workforce data into HDFS for analysis.
  • Hands on experience on AWS cloud services ( Amazon Redshift and Data Pipeline).
  • Extensive Python scripting experience for Scheduling and Process Automation.
  • Experience in Hadoop ecosystem experience in ingestion, storage, querying, processing and analysis of big data.
  • Experience in performing analytics on structured data using Hive queries, operations
  • Experience in implementing Real - Time streaming and analytics using various technologies i.e. Spark streaming and Kafka.
  • Strong experience in database design, writing complex SQL queries and stored procedures using PL/SQL.
  • Experience with Oozie Scheduler in setting up workflow jobs with Map/Reduce and Pig jobs
  • Knowledge and experience of architecture and functionality of NOSQL DB like HBase and Cassandra.
  • Excellent knowledge of waterfall and spiral methodologies of Software Development Life Cycle (SDLC).
  • Excellent technical and analytical skills with clear understanding of design goals of ER modeling for OLTP and dimension modeling for OLAP.
  • Knowledge of Star Schema Modeling, and Snowflake modeling, FACT and Dimensions tables, physical and logical modeling.
  • Experience with data transformations utilizing SnowSQL in Snowflake.
  • Sound knowledge in Data Analysis, Data Validation, Data Cleansing, Data Verification and identifying data mismatch.
  • Expert in generating on-demand and scheduled reports for business analysis or management decision using POWER BI.
  • Experience with Tableau in analysis and creation of dashboard and user stories.
  • Strong experience in using MS Excel and MS Access to dump the data and analyze based on business needs.
  • Experience in writing and executing unit, system, integration and UAT scripts in adata warehouse projects.
  • Ability to learn and adapt quickly to the emerging new technologies.
  • Good communication skills, work ethics and the ability to work in a team efficiently with good leadership skills.

TECHNICAL SKILLS

Big Data: Hadoop3.3, MapReduce, HBase, Pig, Hive, Flume, Sqoop, Spark, Pig, Hive.

Data Modeling Tools: Erwin R2Sp2/9.8/9.7, ER/Studio, Power Designer

ETL Tools: Informatica 10.1/9.6.1, (Power Center), Talend.

Reporting Tools: SSRS, Power BI, Tableau, MS-Excel.

Cloud Services: EC2, S3 RDS, Redshift, Azure Data Lake, Azure Data Factory, GCP and BigQuery.

Databases: Oracle 12c/11g, Teradata R15/R14, MS SQL Server, DB2.

Languages: SQL, PL/SQL, Shell scripting, Unix Shell Script.

Methodologies: JAD, System Development Life Cycle (SDLC), Agile, Waterfall Model

Operating System: Windows, UNIX, Linux

PROFESSIONAL EXPERIENCE

Confidential

Senior Data Engineer

Responsibilities:

  • Working as a Sr. Data Engineer, Assisted in leading the plan, building, and running states within the Enterprise Analytics Team.
  • Involved in all phases of SDLC using Agile and participated in daily scrum meetings with cross teams.
  • Developed and maintained innovative Azure solutions.
  • Analyzed the data flow from different sources to target to provide the corresponding design Architecture in Azure environment.
  • Worked with Azure Data Lake, Azure Data Factory, Databricks, Synapse (SQL Data Warehouse), Azure Blob storage, Azure Storage Explorer.
  • Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big Data technologies.
  • Designed and developed Spark job with Python to implement end to end data pipeline for batch processing.
  • Involved in creating pipeline jobs, scheduling triggers, Mapping data flows using Azure Data Factory(V2) and using Key Vaults to store credentials.
  • Integrated NoSQL database like Hbase with Map Reduce to move bulk amount of data into HBase.
  • Redesigned the Views in snowflake to increase the performance.
  • Involved in designing data warehouses with Azure Synapse.
  • Createddesignsand process flows on how tostandardize Power BI dashboardsto meet thebusiness requirement.
  • Used Azure Data Factory extensively for ingesting data from disparate source systems.
  • Extensively worked with continuous Integration of application using Jenkins.
  • Involved in the solution architecture and Design for data load and the migration of data to Hadoop.
  • Worked with Azure Databricks, Azure Data Factory and Pyspark.
  • Analyzed massive and highly complex HIVE data sets, performing ad-hoc analysis and data manipulation.
  • Generated JSON files from the JSON models created for Zip Code, Group and Claims using Snowflake DB.
  • Implemented Copy activity, custom Azure Data Factory pipeline activities.
  • Worked on Partitioning, Bucketing, Join optimizations and query optimizations in Hive
  • Used MapReduce programs for data cleaning and transformations and load the output into the Hive tables in different file formats
  • Architected and implemented ETL and data movement solutions using Azure Data Platform services (Azure Data Lake, Azure Data Factory, Databricks, Delta lake).
  • Worked on snow-flaking the Dimensions to remove redundancy.
  • Worked on Oozie workflow engine for job scheduling.
  • Created Sqoop job with incremental load to populate Hive External tables.
  • Developed Spark code using Scala and Spark-SQL for faster testing and data processing.
  • Used Spark SQL to process the huge amount of structured data.
  • Developed MapReduce and Pig scripts to cleanse, transform the raw data into meaningful business information and uploaded it into Hive.
  • Worked on loading data into SnowflakeDB in the cloud from various sources.
  • Designed, configured and managed the backup and disaster recovery for HDFS data.
  • UsedAzurereporting services to upload and download reports
  • Implemented Apache Drill on Hadoop to join data from SQL and No SQL databases and store it in Hadoop.
  • Used Git for version control, JIRA for project tracking.
  • Created and Maintained Tables and views in Snowflake.
  • Involved in testing the XML files and checked whether data is parsed and loaded to staging tables.
  • Involved in writing T-SQL programming to implement Stored Procedures.

Confidential - Dayton, OH

Sr. Data Engineer

Responsibilities:

  • As a Data Engineer worked with the analysis teams and management teams and supported them based on their requirements.
  • Followed the Agile methodology to implement the application.
  • Worked with client teams to design and implement modern, scalable data solutions using a range of new and emerging technologies from the Google Cloud Platform.
  • Implemented the Big Data solution using Hadoop, hive to pull/load the data into the HDFS system.
  • Developed and deployed the outcome using spark and Scala code in Hadoop cluster running on GCP.
  • Designed and architected various layer of Data lake.
  • Designed star schema in BigQuery.
  • Build data ingestion from various source systems to Hadoop or GCP using Sqoop, Spark Streaming etc.
  • Worked in GCP Dataproc, GCS, Cloud functions and BigQuery.
  • Extracted data using Sqoop Import query from multiple databases and ingest into Hive tables.
  • Created python scripts to ingest data from on-premise to GCS and built data pipelines using DataFlow for data transformation from GCS to Bigquery.
  • Worked with Google function for event driven data ingestion and route data to Pub sub.
  • Used apache airflow in GCP composer environment to build data pipelines.
  • Installed and configured HDFS, PIG, HIVE, Hadoop and MapReduce.
  • Worked in GCP based Big Data deployments (Real-Time) leveraging BigQuery, Big Table, Google Cloud Storage, Pub Sub, Data Fusion, Dataflow, Dataproc, Airflow, etc.
  • Imported data from RDBMS to HDFS and Hive using Sqoop on regular basis.
  • Monitored Bigquery, Dataproc and cloud Dataflow jobs via Stackdriver for all environments.
  • Involved in debugging and Tuning the PL/SQL code, tuning queries, optimization for the SQL database.
  • Extracted files from MongoDB through Sqoop and placed in HDFS and processed.
  • Worked on POC to check various cloud offerings including Google Cloud Platform (GCP).
  • Developed Source to Target Matrix with ETL transformationlogic for ETL team.
  • Involved in migration of data from existing RDBMS (oracle) to Hadoop using Sqoop for processing data.
  • Worked on google cloud platform (GCP) services like compute engine, cloud load balancing, cloud storage, cloud SQL, stack driver monitoring and cloud deployment manager.
  • Installed and configured Hadoop and responsible for maintaining cluster and managing and reviewing Hadoop log files.
  • Designed and built GCP data driven solutions for enterprise data warehouse and data lakes.
  • Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting on the dashboard.
  • Worked with GCP platform development tools Pub sub, cloud storage, big table, bigquery, dataproc, and composer.
  • Used GCP Console, monitor dataproc cluster and jobs.
  • Developed predictive analytic product by using Apache Spark,SQL.
  • Connected Tableau server to publish dashboard to a central location.
  • Created HiveExternaltables to stage data and then move the data from Staging to main tables
  • Created SSISpackages to migrate data from heterogeneous sources such as MS Excel.
  • Involved in Oozie and workflow scheduler to manage hadoopjobs with control flows.
  • Written T-SQL queries, created dynamic Stored Procedures by using input values.

Confidential - Watertown, MA

Data Engineer

Responsibilities:

  • Worked as a Data Engineer designed and deployed scalable, highly available, and fault tolerant systems on Azure.
  • Utilized SDLC and Agile methodologies such as SCRUM.
  • Provided suggestion to implement multitasking for existing Hive Architecture in Hadoop.
  • Applied data warehousing methodologies in enhancing the existing data model.
  • Designed and implemented database solutions in Azure Data Lake, Azure Data Factory, Azure Synapse Analytics
  • Independently coded new programs and design Tables to load and test the program effectively for the given POC's using Big Data/Hadoop.
  • Developed Python scripts to clean the raw data.
  • Involved in developing PySpark to transform data from one Data Lake to other Data Lake.
  • Effectively worked in Azure Synapse Analytics, Azure Data Lake, Data Factory, Key vault, and Azure Data Bricks.
  • Wrote Hive queries for data analysis to meet the business requirements.
  • Worked with data investigation, discovery and mapping tools to scan every single data record from many sources.
  • Automated various data extraction, transformation, and loading tasks with Python
  • Identified the Entities, attributes and designed a relational database system.
  • Analyzed and designed the business rules for data cleansing that are required by the staging and OLAP & OLTP database.
  • Implemented Star Schema, snowflake methodologies in enhancing data warehouse.
  • Developed SQL Queries to fetch complex data from different tables in remote databases using joins, database links and Bulk collects.
  • Created the ETL data mapping documents between source systems and the target data warehouse.
  • Developed SQL and PL/SQL scripts to transfer tables across the schemas and databases.
  • Configured Input & Output bindings of Azure Function with Azure Cosmos DB collection to read and write data from the container whenever the function executes.
  • Designed solution, developing code and sustainment using T-SQL and Oracle.
  • Worked on creating DDL, DML scripts for the data models.
  • Developed normalized Logical and Physical database models to design OLTP system.
  • Facilitated in developing testing procedures, test cases and User Acceptance Testing (UAT).
  • Designed and developed the data dictionary and Meta data of the models and maintain them.
  • Designed and visualized interactive results using tableau to publish dashboards.

Confidential - Charlotte, NC

Data Modeler

Responsibilities:

  • Massively involved in Data Modeler role to review business requirement and compose source to target data mapping documents.
  • Successfully managed projects using Agile development methodology.
  • Prepared Data Architecture & Data design to present it to client.
  • Participated in JAD sessions for design optimizations related to data structures as well as ETL processes
  • Responsible for defining the naming standards for data warehouse.
  • Worked with MDM systems team with respect to technical aspects and generating reports.
  • Reviewed requirements and designed data model for the Data warehouse to be developed.
  • Created dimensional model for the reporting system by identifying required dimensions and facts using Erwin.
  • Involved in modeling (Star Schema methodologies) in building and designing the logical data model into Dimensional Models.
  • Developed Data mapping, Data Governance, Transformation and Cleansing rules for the Data Management involving OLTP, ODS and OLAP.
  • Created data model design specifications and Source to Target Mapping (STTM) documentation.
  • Applied Data Governance rules for primary qualifier, Class words and valid abbreviation in table name and Column names.
  • Designed data models for MDM system.
  • Generated complex SQL subqueries with inner, outer joins, and aggregate functions to update and delete data in Oracle database.
  • Extensively used normalization techniques (up to 3NF).
  • Developed various operational Drill-through and Drill-down reports using SSRS.
  • Designed and maintained data hub ODS and data marts in data warehouse for reporting.
  • Build the Logical and Physical data model for snowflake as per the changes required.
  • Developed SQL scripts and wrote stored procedures, triggers, and cursors.
  • Performed Data Profiling to understand irregularity and issues, Data mapping, Transformation from Source to Target Database.
  • Developed T-SQL queries, SSIS Packages and Stored Procedures.
  • Developed Stored Procedures, Functions, Packages using PL/SQL.
  • Created queries using BI Reporting variables, navigational attributes and Filters.
  • SQL Query performance tuning to identify and understand performance of the database tables.
  • Involved in creating multiple kinds of Report in Power BI and presented them using Story Points.

We'd love your feedback!