Senior Data Engineer Resume
SUMMARY
- Around 9 years of IT experience in Data Engineering and Data Modeling with high proficiency in using Big Data technologies and cloud infrastructure.
- Good understanding and knowledge with Agile and Waterfall environments.
- Excellent knowledge in Migrating servers, databases, and applications from on premise to Azure and Google Cloud Platform.
- Extensive knowledge of Big Data, Hadoop, MapReduce, Hive and other emerging technologies.
- Developing data pipeline using Sqoop, and MapReduce to ingest workforce data into HDFS for analysis.
- Hands on experience on AWS cloud services ( Amazon Redshift and Data Pipeline).
- Extensive Python scripting experience for Scheduling and Process Automation.
- Experience in Hadoop ecosystem experience in ingestion, storage, querying, processing and analysis of big data.
- Experience in performing analytics on structured data using Hive queries, operations
- Experience in implementing Real - Time streaming and analytics using various technologies i.e. Spark streaming and Kafka.
- Strong experience in database design, writing complex SQL queries and stored procedures using PL/SQL.
- Experience with Oozie Scheduler in setting up workflow jobs with Map/Reduce and Pig jobs
- Knowledge and experience of architecture and functionality of NOSQL DB like HBase and Cassandra.
- Excellent knowledge of waterfall and spiral methodologies of Software Development Life Cycle (SDLC).
- Excellent technical and analytical skills with clear understanding of design goals of ER modeling for OLTP and dimension modeling for OLAP.
- Knowledge of Star Schema Modeling, and Snowflake modeling, FACT and Dimensions tables, physical and logical modeling.
- Experience with data transformations utilizing SnowSQL in Snowflake.
- Sound knowledge in Data Analysis, Data Validation, Data Cleansing, Data Verification and identifying data mismatch.
- Expert in generating on-demand and scheduled reports for business analysis or management decision using POWER BI.
- Experience with Tableau in analysis and creation of dashboard and user stories.
- Strong experience in using MS Excel and MS Access to dump the data and analyze based on business needs.
- Experience in writing and executing unit, system, integration and UAT scripts in adata warehouse projects.
- Ability to learn and adapt quickly to the emerging new technologies.
- Good communication skills, work ethics and the ability to work in a team efficiently with good leadership skills.
TECHNICAL SKILLS
Big Data: Hadoop3.3, MapReduce, HBase, Pig, Hive, Flume, Sqoop, Spark, Pig, Hive.
Data Modeling Tools: Erwin R2Sp2/9.8/9.7, ER/Studio, Power Designer
ETL Tools: Informatica 10.1/9.6.1, (Power Center), Talend.
Reporting Tools: SSRS, Power BI, Tableau, MS-Excel.
Cloud Services: EC2, S3 RDS, Redshift, Azure Data Lake, Azure Data Factory, GCP and BigQuery.
Databases: Oracle 12c/11g, Teradata R15/R14, MS SQL Server, DB2.
Languages: SQL, PL/SQL, Shell scripting, Unix Shell Script.
Methodologies: JAD, System Development Life Cycle (SDLC), Agile, Waterfall Model
Operating System: Windows, UNIX, Linux
PROFESSIONAL EXPERIENCE
Confidential
Senior Data Engineer
Responsibilities:
- Working as a Sr. Data Engineer, Assisted in leading the plan, building, and running states within the Enterprise Analytics Team.
- Involved in all phases of SDLC using Agile and participated in daily scrum meetings with cross teams.
- Developed and maintained innovative Azure solutions.
- Analyzed the data flow from different sources to target to provide the corresponding design Architecture in Azure environment.
- Worked with Azure Data Lake, Azure Data Factory, Databricks, Synapse (SQL Data Warehouse), Azure Blob storage, Azure Storage Explorer.
- Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big Data technologies.
- Designed and developed Spark job with Python to implement end to end data pipeline for batch processing.
- Involved in creating pipeline jobs, scheduling triggers, Mapping data flows using Azure Data Factory(V2) and using Key Vaults to store credentials.
- Integrated NoSQL database like Hbase with Map Reduce to move bulk amount of data into HBase.
- Redesigned the Views in snowflake to increase the performance.
- Involved in designing data warehouses with Azure Synapse.
- Createddesignsand process flows on how tostandardize Power BI dashboardsto meet thebusiness requirement.
- Used Azure Data Factory extensively for ingesting data from disparate source systems.
- Extensively worked with continuous Integration of application using Jenkins.
- Involved in the solution architecture and Design for data load and the migration of data to Hadoop.
- Worked with Azure Databricks, Azure Data Factory and Pyspark.
- Analyzed massive and highly complex HIVE data sets, performing ad-hoc analysis and data manipulation.
- Generated JSON files from the JSON models created for Zip Code, Group and Claims using Snowflake DB.
- Implemented Copy activity, custom Azure Data Factory pipeline activities.
- Worked on Partitioning, Bucketing, Join optimizations and query optimizations in Hive
- Used MapReduce programs for data cleaning and transformations and load the output into the Hive tables in different file formats
- Architected and implemented ETL and data movement solutions using Azure Data Platform services (Azure Data Lake, Azure Data Factory, Databricks, Delta lake).
- Worked on snow-flaking the Dimensions to remove redundancy.
- Worked on Oozie workflow engine for job scheduling.
- Created Sqoop job with incremental load to populate Hive External tables.
- Developed Spark code using Scala and Spark-SQL for faster testing and data processing.
- Used Spark SQL to process the huge amount of structured data.
- Developed MapReduce and Pig scripts to cleanse, transform the raw data into meaningful business information and uploaded it into Hive.
- Worked on loading data into SnowflakeDB in the cloud from various sources.
- Designed, configured and managed the backup and disaster recovery for HDFS data.
- UsedAzurereporting services to upload and download reports
- Implemented Apache Drill on Hadoop to join data from SQL and No SQL databases and store it in Hadoop.
- Used Git for version control, JIRA for project tracking.
- Created and Maintained Tables and views in Snowflake.
- Involved in testing the XML files and checked whether data is parsed and loaded to staging tables.
- Involved in writing T-SQL programming to implement Stored Procedures.
Confidential - Dayton, OH
Sr. Data Engineer
Responsibilities:
- As a Data Engineer worked with the analysis teams and management teams and supported them based on their requirements.
- Followed the Agile methodology to implement the application.
- Worked with client teams to design and implement modern, scalable data solutions using a range of new and emerging technologies from the Google Cloud Platform.
- Implemented the Big Data solution using Hadoop, hive to pull/load the data into the HDFS system.
- Developed and deployed the outcome using spark and Scala code in Hadoop cluster running on GCP.
- Designed and architected various layer of Data lake.
- Designed star schema in BigQuery.
- Build data ingestion from various source systems to Hadoop or GCP using Sqoop, Spark Streaming etc.
- Worked in GCP Dataproc, GCS, Cloud functions and BigQuery.
- Extracted data using Sqoop Import query from multiple databases and ingest into Hive tables.
- Created python scripts to ingest data from on-premise to GCS and built data pipelines using DataFlow for data transformation from GCS to Bigquery.
- Worked with Google function for event driven data ingestion and route data to Pub sub.
- Used apache airflow in GCP composer environment to build data pipelines.
- Installed and configured HDFS, PIG, HIVE, Hadoop and MapReduce.
- Worked in GCP based Big Data deployments (Real-Time) leveraging BigQuery, Big Table, Google Cloud Storage, Pub Sub, Data Fusion, Dataflow, Dataproc, Airflow, etc.
- Imported data from RDBMS to HDFS and Hive using Sqoop on regular basis.
- Monitored Bigquery, Dataproc and cloud Dataflow jobs via Stackdriver for all environments.
- Involved in debugging and Tuning the PL/SQL code, tuning queries, optimization for the SQL database.
- Extracted files from MongoDB through Sqoop and placed in HDFS and processed.
- Worked on POC to check various cloud offerings including Google Cloud Platform (GCP).
- Developed Source to Target Matrix with ETL transformationlogic for ETL team.
- Involved in migration of data from existing RDBMS (oracle) to Hadoop using Sqoop for processing data.
- Worked on google cloud platform (GCP) services like compute engine, cloud load balancing, cloud storage, cloud SQL, stack driver monitoring and cloud deployment manager.
- Installed and configured Hadoop and responsible for maintaining cluster and managing and reviewing Hadoop log files.
- Designed and built GCP data driven solutions for enterprise data warehouse and data lakes.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting on the dashboard.
- Worked with GCP platform development tools Pub sub, cloud storage, big table, bigquery, dataproc, and composer.
- Used GCP Console, monitor dataproc cluster and jobs.
- Developed predictive analytic product by using Apache Spark,SQL.
- Connected Tableau server to publish dashboard to a central location.
- Created HiveExternaltables to stage data and then move the data from Staging to main tables
- Created SSISpackages to migrate data from heterogeneous sources such as MS Excel.
- Involved in Oozie and workflow scheduler to manage hadoopjobs with control flows.
- Written T-SQL queries, created dynamic Stored Procedures by using input values.
Confidential - Watertown, MA
Data Engineer
Responsibilities:
- Worked as a Data Engineer designed and deployed scalable, highly available, and fault tolerant systems on Azure.
- Utilized SDLC and Agile methodologies such as SCRUM.
- Provided suggestion to implement multitasking for existing Hive Architecture in Hadoop.
- Applied data warehousing methodologies in enhancing the existing data model.
- Designed and implemented database solutions in Azure Data Lake, Azure Data Factory, Azure Synapse Analytics
- Independently coded new programs and design Tables to load and test the program effectively for the given POC's using Big Data/Hadoop.
- Developed Python scripts to clean the raw data.
- Involved in developing PySpark to transform data from one Data Lake to other Data Lake.
- Effectively worked in Azure Synapse Analytics, Azure Data Lake, Data Factory, Key vault, and Azure Data Bricks.
- Wrote Hive queries for data analysis to meet the business requirements.
- Worked with data investigation, discovery and mapping tools to scan every single data record from many sources.
- Automated various data extraction, transformation, and loading tasks with Python
- Identified the Entities, attributes and designed a relational database system.
- Analyzed and designed the business rules for data cleansing that are required by the staging and OLAP & OLTP database.
- Implemented Star Schema, snowflake methodologies in enhancing data warehouse.
- Developed SQL Queries to fetch complex data from different tables in remote databases using joins, database links and Bulk collects.
- Created the ETL data mapping documents between source systems and the target data warehouse.
- Developed SQL and PL/SQL scripts to transfer tables across the schemas and databases.
- Configured Input & Output bindings of Azure Function with Azure Cosmos DB collection to read and write data from the container whenever the function executes.
- Designed solution, developing code and sustainment using T-SQL and Oracle.
- Worked on creating DDL, DML scripts for the data models.
- Developed normalized Logical and Physical database models to design OLTP system.
- Facilitated in developing testing procedures, test cases and User Acceptance Testing (UAT).
- Designed and developed the data dictionary and Meta data of the models and maintain them.
- Designed and visualized interactive results using tableau to publish dashboards.
Confidential - Charlotte, NC
Data Modeler
Responsibilities:
- Massively involved in Data Modeler role to review business requirement and compose source to target data mapping documents.
- Successfully managed projects using Agile development methodology.
- Prepared Data Architecture & Data design to present it to client.
- Participated in JAD sessions for design optimizations related to data structures as well as ETL processes
- Responsible for defining the naming standards for data warehouse.
- Worked with MDM systems team with respect to technical aspects and generating reports.
- Reviewed requirements and designed data model for the Data warehouse to be developed.
- Created dimensional model for the reporting system by identifying required dimensions and facts using Erwin.
- Involved in modeling (Star Schema methodologies) in building and designing the logical data model into Dimensional Models.
- Developed Data mapping, Data Governance, Transformation and Cleansing rules for the Data Management involving OLTP, ODS and OLAP.
- Created data model design specifications and Source to Target Mapping (STTM) documentation.
- Applied Data Governance rules for primary qualifier, Class words and valid abbreviation in table name and Column names.
- Designed data models for MDM system.
- Generated complex SQL subqueries with inner, outer joins, and aggregate functions to update and delete data in Oracle database.
- Extensively used normalization techniques (up to 3NF).
- Developed various operational Drill-through and Drill-down reports using SSRS.
- Designed and maintained data hub ODS and data marts in data warehouse for reporting.
- Build the Logical and Physical data model for snowflake as per the changes required.
- Developed SQL scripts and wrote stored procedures, triggers, and cursors.
- Performed Data Profiling to understand irregularity and issues, Data mapping, Transformation from Source to Target Database.
- Developed T-SQL queries, SSIS Packages and Stored Procedures.
- Developed Stored Procedures, Functions, Packages using PL/SQL.
- Created queries using BI Reporting variables, navigational attributes and Filters.
- SQL Query performance tuning to identify and understand performance of the database tables.
- Involved in creating multiple kinds of Report in Power BI and presented them using Story Points.
