We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

Rochester, MN

SUMMARY

  • A qualified IT Professional with 8+ years of experience in Data Analysis, Data Warehouse Concepts and Hadoop ecosystem. Good technical expertise in SQL, Python scripting, Hadoop technologies and AWS Cloud Services.
  • Used Business Intelligence tools such as Business Objects and Data Visualization tools such as Tableau, Power BI.
  • Domain Knowledge of Finance, Logistics and Health insurance.
  • Worked with several Azure services, such as Data Lake, to store and analyze data.
  • Experience designing business intelligence solutions using Microsoft SQL Server 2008 and 2012.
  • Mastery of DBMS concepts like PL/SQL, SQL, Oracle.
  • Created an Azure SQL database, monitored it, and restored it. Migrated Microsoft SQL server to Azure SQL database.
  • Experience with Azure Cloud, Azure Data Factory, Azure Data Lake Storage, Azure Synapse Analytics, Azure Analytical services, Big Data Technologies (Apache Spark), and Data Bricks is preferred.
  • Extensive experience developing and implementing cloud architecture on Microsoft Azure.
  • Extensive expertise with Amazon Web Services such as Amazon EC2, S3, RDS, IAM, Auto Scaling, CloudWatch, SNS, Athena, Glue, Kinesis, Lambda, EMR, Redshift, and DynamoDB.
  • Created a connection from Azure to an on - premises data center using the Azure Express Route for Single and Multi-Subscription.
  • Worked on ETL Migration services by creating and deploying AWS Lambda functions to provide a serverless data pipeline that can be written to Glue Catalog and queried from Athena.
  • Developed ETL pipelines in and out of the data warehouse using a mix of Python and Snowflakes SnowSQL Writing SQL queries against Snowflake.
  • Analytics and cloud migration from on-premises to AWS Cloud with AWS EMR, S3, and DynamoDB.
  • Experience in creating and managing reporting and analytics infrastructure for internal business clients using AWS services including Athena, Redshift, Spectrum, EMR, and Quick Sight.
  • Strong Experience in Data Engineering, Data Pipeline Design, Development, Documentation, Deployment and Integration as a Sr. Data Engineer/Data Developer and Data Modeler.
  • Experience in layers of Hadoop Framework - Storage (HDFS), Analysis (Pig and Hive), Engineering (Jobs and Workflows), extending the functionality by writing custom UDFs.
  • Strong experience on Hadoop distributions like Horton works and Cloudera.
  • Hands-on experience with Hadoop architecture and various components such as Hadoop File System HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Hadoop MapReduce programming.
  • Extensive knowledge of data architecture including designing pipelines, data ingestion, Hadoop/Spark architecture, data modeling, data mining, machine learning and advanced data processing.
  • Extensive experience in developing Data warehouse applications using Hadoop, Informatica, Oracle, Teradata, MS SQL server on UNIX and Windows platforms and experience in creating complex mappings using various transformations and developing strategies for Extraction, Transformation and Loading (ETL) mechanism by using Informatica.
  • Experience designing error and exception handling techniques for detecting, recording, and reporting errors.
  • Extensive experience in T-SQL writing stored procedures, triggers, functions, tables, views, indexes, and relational database models.
  • Experience in writing Unit Test and Smoke Test for testing the code (modules) using Scala Test Framework.
  • Experience in Installing, Upgrading and Configuring Microsoft SQL Server
  • Solid programming knowledge on Python, PySpark, SQL and shell scripting.
  • Extract, transform and load the data from different formats like JSON, a Database, and expose it for ad-hoc/interactive queries using Spark SQL.
  • Good working experience in Relational and databases like MySQL, Oracle.
  • Developed framework for converting existing PowerCenter mappings and to PySpark(Python and Spark) Jobs.
  • Create Pyspark frame to bring data from DB2 to Amazon S3.
  • Worked on developing PySpark script to encrypting the raw data by using hashing algorithms concepts on client specified columns.
  • Experience in writing SQL queries for creating tables, views and applying filters in snowflake.
  • Familiar with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining, machine learning and advanced data processing.
  • Experience optimizing ETL workflows.
  • Day to-day responsibility includes developing ETL Pipelines in and out of data warehouse, develop major regulatory and financial reports using advanced SQL queries in snowflake.
  • Implement One Time Data Migration of Multistate level data from SQL server to Snowflake by using Python and SnowSQL.
  • Excellent knowledge in Data Analysis, Data Validation, Data Cleansing, Data Verification and identifying data mismatch.

PROFESSIONAL EXPERIENCE

Data Engineer

Confidential, Rochester MN

Responsibilities:

  • Experience in using and tuning relational databases (e.g., Microsoft SQL Server, Oracle, MySQL) and columnar databases (e.g., Amazon Redshift, Microsoft SQL Data Warehouse)
  • Extensive expertise using the core Spark APIs and processing data on a EMR cluster
  • Programmed in Python, PySpark and SQL to streamline the incoming data and build the data pipelines to get the useful insights, and orchestrated pipelines
  • Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
  • Used Snowflake data warehouse to perform data analysis by applying filters on top of tables and views
  • Created multiple notebooks in Databricks to perform various transformations as per the business requirements
  • Worked on migrating data from on-premises Bigdata platform to AWS cloud architecture.
  • Developed Scala scripts using both Data frames/SQL/Data sets and RDD/MapReduce in Spark for Data Aggregation, queries and writing data back into OLTP system through Sqoop
  • Supporting Continuous storage in AWS using Elastic Block Storage, S3, Glacier. Created Volumes and configured Snapshots for EC2 instances
  • Developed Hive queries to pre-process the data required for running the business process
  • Worked on root cause analysis for the all the issues that occur in production or batch and provide the resolution for the issues.
  • Worked on ETL Migration services by developing and deploying AWS Lambda functions for generating a serverless data pipeline which can be written to Glue Catalog and can be queried from Athena.
  • Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.
  • Expert in working with Hive data warehouse tool-creating tables, data distribution by implementing partitioning and bucketing, writing, and optimizing the HiveQL queries.
  • Extensive Shell/Python scripting experience and setup production systems on UNIX/LINUX Environment, Comprehensive experience in developing simple to complex Map reduce and Streaming jobs using Scala and Java for data cleansing, filtering, and data aggregation. Also possess detailed knowledge of MapReduce framework.
  • Involved in Designing and Developing Enhancements of CSG using AWS APIS
  • Supporting Continuous storage in AWS using Elastic Block Storage, S3, Glacier. Created Volumes and configured Snapshots for EC2 instances.
  • Built PHP applications to meet product requirements and satisfy use cases using MVC architecture, CodeIgniter Framework and Drupal CMS
  • Done installing, configuring, and maintaining Code Igniter, PHP, Apache, and MySQL on AWS Cloud Servers
  • Involved in the Development of various layers to accommodate the application as per the MVC design pattern, DAO and DTO patterns using Spring and struts and hibernate.

Environment: AWS, PySpark, Hadoop, HDFS, ADF, Databricks, Pig, Sqoop, Hive, NoSQL, HBase, Shell Scripting, Python, Scala, Spark, Spark SQL, AWS, SQL Server, Tableau, ETL, PHP.

Data Engineer

Confidential, MI

Responsibilities:

  • Writing MapReduce code using python to get rid of certain security issues in the data.
  • Synchronizing both the unstructured and structured data using Hive on business prospectus.
  • Used Pig Latin at client-side cluster and HiveQL at server-side cluster.
  • Importing the complete data from RDBMS to HDFS cluster using Sqoop.
  • Creating external tables and moving the data onto the tables from managed tables.
  • Moving this partitioned data onto the different tables as per as business requirements.
  • Partitioning and bucketing the imported data using HiveQL.
  • Partitioning dynamically using dynamic partition insert feature.
  • Setting up the work schedule using Airflow.
  • Identifying the errors in the logs and rescheduling/resuming the job
  • Involved in Designing the SRS with Activity Flow Diagrams using UML.
  • Worked with data transfer from on-premise SQL servers to cloud databases (Azure Synapse Analytics (DW) and Azure SQL DB).
  • Created Pipelines which were built in Azure Data Factory using Linked Services/Datasets/Pipeline/ to extract, transform, and load data from a variety of sources including Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and reverse.
  • Created CI-CD Pipelines using Azure DevOps
  • Used a blend of Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics, gather, convert, and load the data from source systems to Azure Data Storage services.
  • Employed Agile methodology for project management, including tracking project milestones; gathering project requirements and technical closures; planning and estimation of project effort; creating important project related design documents and identifying technology related risks and issues.
  • Used IDEs like Eclipse, IntelliJ IDE, PyCharm IDE, Notepad ++, and Visual Studio for development.
  • Designed the Data Marts in dimensional data modelling using star and snowflake schemas.
  • Experience in writing SQL queries against snowflake.
  • Excellent Software Development Life Cycle (SDLC) with good working knowledge of testing methodologies, disciplines, tasks, resources, and scheduling.
  • Excellent knowledge in Data Analysis, Data Validation, Data Cleansing, Data Verification and identifying data mismatch.
  • Experience in designing error and exception handling procedures to identify, record and report errors.
  • Create Migration request for IBM InfoSphere DataStage, SQL, and AutoSys (Jils) from development to test.
  • Scheduling the IBM InfoSphere DataStage job using AutoSys.
  • Analyze the request from business & other teams and do research and generate the PLSQL and T-SQL queries and provide the data as needed
  • Good hands-on experience on Loading data onto Hive from Spark RDD’s.
  • Worked on Spark SQL UDF’s and Hive UDF’s.
  • Experience on Spark with Scala/Python.
  • Working on Statefull Transformations in Spark Streaming.
  • Worked on Batch processing and Real-time data processing on Spark Streaming using Lambda architecture.

Environment: Python, Mocrosoft Azure, Airflow, Hortonworks distribution, IBM Data Stage, Map Reduce, HDFS, Spark, Scala, Python Hive, HBase, SQL, Sqoop, Flume, Oozie, Apache, Tez, Tableau. Tek Leaders

Big Data Developer

Confidential, Columbus, Indiana

Responsibilities:

  • Developed tables, packages, procedures, functions, collections, triggers, cursors, ref cursors, exceptions, views, synonyms, performance tuning, interfaces, API, Lookups, and processing constraints.
  • Involved in documentation of functional and technical requirements specification.
  • Worked on preparation of estimation plan to implement the change request based on the code freeze dates in different instances
  • Involved in the development of the new AWS Fargate API, which is comparable to the ECS run task API.
  • Worked on the code transfer of a quality monitoring application from AWS EC2 to AWS Lambda, as well as the construction of logical datasets to administer quality monitoring on snowflake warehouses.
  • Comprehensive teamwork with the client to gather requirements for solutions
  • Completed requirement analysis and compiled a list of clarifications and issues.
  • Designed, developed, and performed maintenance of data integration in Hadoop and RDBMS environment with both traditional and non-traditional source system as well as RDBMS and NoSQL data stores for data access and analysis.
  • Involved in recovery of Hadoop clusters and worked on cluster size of 310 nodes.
  • Worked on creating Hive tables, loading, and analyzing data using Hive queries.
  • Experience in proving application support for Jenkins.
  • Developed a data pipeline with AWS to extract the data from weblogs and store in HDFS.
  • Used Hive QL to analyze the partitioned and bucketed data and compute various metrics for reporting.
  • Used reporting tools like Power BI to connect with Hive for generating daily reports of data.
  • Responsible for day-to-day Production Support operations, Job monitoring, Incident ticket resolution, on time delivers and code deployment
  • Developed DE fix scripts for hold orders and corrected the process for old existing orders.

Environment: Python, Bigdata, Hadoop, HBase, Hive, Spark, PySpark, ClouderaKafka, Sqoop, Jenkins, Unix Shell scripting, GitHub, SQL, Power BI.

Hadoop Developer

Confidential

Responsibilities:

  • Designed and Developed data integration/engineering workflows on big data technologies and platforms like Hadoop, Spark, MapReduce, Hive and HBase.
  • Worked in Agile methodology and actively participated in standup calls, PI planning.
  • Involved in Requirement gathering and prepared the Design documents.
  • Involved in importing data into HDFS and Hive using Sqoop, in involved in creating Hive tables, loading with data, and writing Hive queries.
  • Developed Hive queries and Sqooped data from RDBMS to Hadoop staging area.
  • Handled importing of data from various data sources, performed transformations using Hive, and loaded data into data lake.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities.
  • Processed data stored in data lake, created external tables using Hive and developed scripts to ingest and repair tables that can be reused across the project.
  • Developed dataflows and processes for Data processing using SQL (SparkSQL & Dataframes).
  • Designed and developed Map Reduce (hive) programs to analyze and evaluate multiple solutions by considering multiple cost factors across the business as well operational impact.
  • Involved in planning process of iterations under the Agile Scrum methodology.
  • Working on Hive Metastore backup, Partitioning, and bucketing techniques in hive to improve the performance.
  • Scheduling Spark jobs using Oozie workflow in Hadoop Cluster and Generated detailed design documentation for the source-to-target transformations.

Environment: Spark, Python, Sqoop, Hive, Hadoop, SQL, HBase, MapReduce, HDFS, Oozie, Agile.

SQL Developer

Confidential

Responsibilities:

  • Involved in designing and managing schema objects such as Tables, Views, Indexes, Stored Procedures, Triggers and maintained referential integrity using SQL Server Management Studio.
  • Processed ETL to transfer data from remote data centers to local data centers using SSIS. Cleansing, messaging of the data is done on the local database.
  • Validated Historical data load from Legacy System into Netezza Datamart.
  • Worked in all the entities like customer records, payment records have been loaded target database according to its own table structure during the migration process.
  • Tested database integrity referential integrity and constrains during the database migration testing process.
  • Provides policy administration and technical services to commercial property/casualty insurers of all sizes.
  • Presented Daily Analysis Report to highlight bugs, issues and risks involved.
  • Constructed and implemented multiple-table links requiring complex join statements, including Outer-Joins and Self-Joins.
  • Developed Shell scripts for job automation and daily backup.
  • Created documentation of the business process through Designer.

Environment: Informatica PowerCenter, Oracle, PL/SQL, MS SQL Server, Powerbi, Cognos, Autosys, and Quality Center.

We'd love your feedback!