We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

Columbus, IndianA

SUMMARY

  • Over 8 years of IT experience in a variety of industries working on Big Data technology using technologies such as Cloudera and Hortonworks distributions. Hadoop working environment includes Hadoop, Data Bricks, Spark, MapReduce, Power BI, Power BI pro, Denodo Data virtualization, Power BI Mobile, Kafka, Hive, Ambari, Sqoop, HBase, Unique combination of technical and Impala.
  • Fluent programming experience with Scala, Python, SQL, T - SQL, R.
  • Hands-on experience in developing and deploying enterprise-based applications using major Hadoop ecosystem components like MapReduce, YARN, Hive, HBase, Flume, Sqoop, Spark MLlib, Power BI, Spark GraphX, Spark SQL, Danodo Data virtualization, Kafka.
  • Used Power BI Desktop to develop data analysis multiple data source to visualize the reports.
  • Adept at configuring and installing Hadoop/Spark Ecosystem Components.
  • Proficient with Spark Core, Spark SQL, Spark MLlib, Spark GraphX and Spark Streaming for processing and transforming complex data using in-memory computing capabilities written in Scala. Worked with Spark to improve efficiency of existing algorithms using Spark Context, Spark SQL, Spark MLlib, Data Frame, Pair RDD's and Spark YARN.
  • Experience in application of various data sources like Oracle SE2, SQL Server, Flat Files and Unstructured files into a data warehouse.
  • Orchestrated all Data Pipelines using Azure Data Factory and Built a custom alert platform for monitoring.
  • Able to use Sqoop to migrate data between RDBMS, NoSQL databases and HDFS.
  • Experience in designing Azure Cloud Architecture and Implementation plans for hosting complex application workloads on MS Azure.
  • Very keen in knowing newer techno stack that Google Cloud platform(GCP)adds.
  • Experience in Extraction, Transformation and Loading (ETL) data from various sources into Data Warehouses, as well as data processing like collecting, aggregating and moving data from various sources using Apache Flume, Kafka, PowerBI and Microsoft SSIS.
  • Collaborated with fellow consultants to brainstorm and develop unique solution for each client's unique implementation.
  • Used Azure Data Factory extensively for ingesting data from disparate source systems.
  • Built a common sftp download or upload frameworks using Azure Data Factory and Databricks.
  • Hands-on experience with Hadoop architecture and various components such as Hadoop File System HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Hadoop MapReduce programming.
  • Used various sources to pull data into Power BI such as SQL Server, SAP BW, Oracle, SQL Azure etc.
  • Comprehensive experience in developing simple to complex Map reduce and Streaming jobs using Scala and Java for data cleansing, filtering and data aggregation. Also possess detailed knowledge of MapReduce framework.
  • Experience in Azure Marketplace where to search, deploy and purchase wide range of applications and services.
  • Used IDEs like Eclipse, IntelliJ IDE, PyCharm IDE, Notepad ++, and Visual Studio for development.
  • Architect and implement ETL and Data movement solution using Azure Data Factory(ADS), SSIS.
  • Responsible for building reports in Power BI from the Scratch.
  • Seasoned practice in Machine Learning algorithms and Predictive Modeling such as Linear Regression, Logistic Regression, Naïve Bayes, Decision Tree, Random Forest, KNN, Neural Networks, and K-means Clustering.
  • Manage Azure Synapse Analytics and Denodo Data Virtualization platform.
  • Ample knowledge of data architecture including data ingestion pipeline design, Hadoop/Spark architecture, data modeling, data mining, machine learning and advanced data processing.
  • Responsible for estimating the cluster size, monitoring and troubleshooting of the Spark data bricks cluster.
  • Experience in managing Azure Storage Accounts.
  • Worked with both live and import data into Power BI using star schema.
  • Implementation Copy activity, Custom Azure Data Factory Pipeline Activities for On-could ETL processing.
  • Experience working with NoSQL databases like Cassandra and HBase and developed real-time read/write access to very large datasets via HBase.
  • Developed Spark Applications that can handle data from various RDBMS (MySQL, Oracle Database) and Streaming sources.
  • Legacy information logical code migrated into python, Denodo Data virtualization, Spark, Data Bricks used for ETL process and loaded Dataset results into cloud storage S3/blob/ relational databases.
  • Proficient SQL experience in querying, data extraction/transformations and developing queries for a wide range of applications.
  • Experience in migration om premise to Windows Azure using Azure Site Recovery and Azure backups.
  • Capable of processing large sets (Gigabytes) of structured, semi-structured or unstructured data.
  • Experience in analyzing data using HiveQL, Pig, HBase and custom MapReduce programs in Java 8.
  • Used different types of slicers available in Power BI for creating reports.
  • Experience working with GitHub/Git 2.12 source and version control systems.
  • Hands-on experience using various AWS Services including EC2, EMR cluster, Redshift, Data Bricks, S3 Buckets, AWS kinesis and Iaas/ Paas/SaaS.

PROFESSIONAL EXPERIENCE

Data Engineer

Confidential, Columbus, Indiana

Responsibilities:

  • Worked on AWS Data pipeline to configure data loads from S3 to into Redshift.
  • Using AWS Redshift, I Extracted, transformed and loaded data from various heterogeneous data sources and destinations
  • Developed algorithm for Fuzzy Matching of vendor records for de-duping data and creating a unique record set.
  • Created Tables, Stored Procedures, and extracted data using T-SQL for business users whenever required.
  • Legacy information batch/ real time ETL logical code migrated into Hadoop using Python, Spark Context level of Parallelism and memory tuning.
  • Performs data analysis and design, and creates and maintains large, complex logical and physical data models, and metadata repositories using ERWIN and MB MDR
  • Worked on creating dependenciesofactivities in Azure Data Factory.
  • Develop Azure POC for marketing Data Research.
  • I have written shell script to trigger data Stage jobs.
  • Data profiling,Data Analyst; Identify and implement business rules to uniquely identify Securities.
  • Assist service developers in finding relevant content in the existing reference models.
  • Provide leadership and oversight regarding design and implementation of Denodo Data virtualization to integrate data from disparate data source and appears to as one uniform source.
  • Creating pipelines with GUlin Azure Data Factory.
  • Like Access, Excel, CSV, Oracle, flat files using connectors, tasks and transformations provided by AWS Data Pipeline.
  • Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.
  • Developed pipelines to move the data from Azure blob storage/ file share to Azure SQL data warehouse and blob.
  • Worked on developing Pyspark script to encrypting the raw data by using Hashing algorithms concepts on client specified columns.
  • Responsible for Design, Development, and testing of the database and Developed Stored Procedures, Views, and Triggers
  • Transforming data in Azure Data Factory with tha ADF Transformations.
  • Developed Python-based API (RESTful Web Service) to track revenue and perform revenue analysis.
  • Compiling and validating data from all departments and Presenting to Director Operation.
  • Design and implement database solution in Azure SQL Data Warehouse, Azure SQL.
  • KPI calculator Sheet and maintain that sheet within SharePoint.
  • Created Tableau reports with complex calculations and worked on Ad-hoc reporting using PowerBI.
  • Knowledge of Data Integration and Virtualization using Denodo platform and Denodo data virtualization certified Developer.
  • Expert in using Databricks with Azure Data Factory(ADF)to compute large volumes of data.
  • Creating data model that correlates all the metrics and gives a valuable output.
  • Worked on the tuning of SQL Queries to bring down run time by working on Indexes and Execution Plan.
  • Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
  • Azure Data Factory(ADF), Integration RUN Time(IR), File System Data Ingestion, Relational Data Ingestion.
  • Pre-processing using Hive and Pig.
  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics.
  • Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in In Azure Databricks.
  • Implemented Copy activity, Custom Azure Data Factory Pipeline Activities
  • Primarily involved in Data Migration using SQL, SQL Azure, Azure Storage, and Azure Data Factory, SSIS, PowerShell.
  • Developed custom alerts using Azure Data Factory, SQLDB and Logic App.
  • Architect & implement medium to large scale BI solutions on Azure using Azure Data Platform services (Azure Data Lake, Data Factory, Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQL DB).
  • Worked in mixed role DevOps: Azure Arcitect /System Engineer. Network operations and Data Engineering.
  • Migration of on-premise data (Oracle/ SQL Server/ DB2/ MongoDB) to Azure Data Lake and Stored (ADLS) using Azure Data Factory (ADF V1/V2).
  • Worked with multiple teams for Business Denodo data virtualization, Database, Flat File, Stored Procedure, Views, Packages, Unix Scripts.
  • Configure and Manage VM Operations such as Snapshot, Clone, FT& Template.
  • VM Tool and VM Hardware upgradations in regular intervals.
  • Virtual Machine hardware up gradation tasks (VM Disk resize, Increase Memory and CPU).
  • Virtual Machine Operations such as Snapshot, Clone, FT and VM Reconfigurations.
  • Developed a detailed project plan and helped manage the data conversion migration from the legacy system to the target snowflake database.
  • Design, develop, and test dimensional data models using Star and Snowflake schema methodologies under the Kimball method.
  • Implement ad-hoc analysis solutions using Azure Data Lake Analytics/Store, HDInsight
  • Developed data pipeline using Spark, Hive, Pig, python, Impala, and HBase to ingest customer
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Ensure deliverables (Daily, Weekly & Monthly MIS Reports) are prepared to satisfy the project requirements cost and schedule
  • Worked on a direct query using PowerBI to compare legacy data with the current data and generated reports and stored and dashboards.
  • Designed SSIS Packages to extract, transfer, load (ETL) existing data into SQL Server from different environments for the SSAS cubes (OLAP)
  • SQL Server reporting services (SSRS). Created & formatted Cross-Tab, Conditional, Drill-down, Top N, Summary, Form, OLAP, Subreports, ad-hoc reports, parameterized reports, interactive reports & custom reports
  • Created action filters, parameters and calculated sets for preparing dashboards and worksheets using PowerBI
  • Developed visualizations and dashboards using PowerBI
  • Used ETL to implement the Slowly Changing Transformation, to maintain Historically Data in Data warehouse.
  • Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
  • Created dashboards for analyzing POS data using Power BI

Environment: MS SQL Server 2016, T-SQL, SQL Server Integration Services (SSIS), Azure Data Lake, Data Factory, SQL Server Reporting Services (SSRS), SQL Server Analysis Services (SSAS), Azure Data bricks, Azure SQL, Management Studio (SSMS), Advance Excel (creating formulas, pivot tables, Hlookup, Vlookup, Macros), Spark, Python, ETL, Data Bricks, Power BI, Tableau, Hive/Hadoop, Azure SQL Data warehouse, Snowflakes, Power BI, AWS Data Pipeline, IBM Cognos 10.1, Data Stage.

Data Engineer

Confidential, Columbus, OH

Responsibilities:

  • Implemented Apache Airflow for authoring, scheduling and monitoring Data Pipelines
  • Solid knowledge of Power BI and Tableau Desktop report performance optimization.
  • Designed several DAGs (Directed Acyclic Graph) for automating ETL pipelines
  • Ingested data in mini-batches and preforms RDD transformations on those mini-batches of data by using Spark Streaming to perform Streaming analytics in Data Bricks.
  • Performed data extraction, transformation, loading, and integration in data warehouse, operational data stores and master data management
  • Strong understanding of AWS components such as EC2 and S3
  • Performed Data Migration to GCP
  • Responsible for data services and data movement infrastructures
  • Experienced in ETL concepts, building ETL solutions and Data modeling
  • Expert on Microsoft Power BI and Tableau reports, dashboards and publishing to the end users for executive level Business Decision.
  • Worked on architecting the ETL transformation layers and writing spark jobs to do the processing.
  • Aggregated daily sales team updates to send report to executives and to organize jobs running on Spark clusters
  • Loaded application analytics data into data warehouse in regular intervals of time
  • Designed & build infrastructure for the Google Cloud environment from scratch
  • Expert on maintaining and managing Tableau and Power BI driven reports and dashboards.
  • Experienced in fact dimensional modeling (Star schema, Snowflake schema), transactional modeling and SCD (Slowly changing dimension)
  • Leveraged cloud and GPU computing technologies for automated machine learning and analytics pipelines, such as AWS, GCP
  • Worked on confluence and Jira
  • Expert in creating multiple kind of Power BI Reports and Dashboards.
  • Designed and implemented configurable data delivery pipeline for scheduled updates to customer facing data stores built with Python
  • Proficient in Machine Learning techniques (Decision Trees, Linear/Logistic Regressors) and Statistical Modeling
  • Compiled data from various sources to perform complex analysis for actionable results
  • Measured Efficiency of Hadoop/Hive environment ensuring SLA is met
  • Optimized the Tensorflow Model for efficiency
  • Expert in creating multiple kinds of Report in Power BI and presented them using Story Points.
  • Analyzed the system for new enhancements/functionalities and perform Impact analysis of the application for implementing ETL changes
  • Implemented a Continuous Delivery pipeline with Docker, and Git Hub and AWS
  • Reporting on VM improvements and made recommendations for the upgrades, which include VM Hadware versions and VMware Tools with less business impact.
  • Managing Tasks, Events and Alarms.
  • Maintain and administer all virtual backups with EMC Avamar/ Data Domain Backup and Recovery technology. Ownership on Storage Migration projects from Old SAN to New SAN.
  • Responsible for creating SQL datasets for Power BI and Ad-hoc Reports.
  • Enabled Monitoring and alerting of virtual environment with the utilization of Solarwinds VM Monitor, vCops & vRops. Deployment of Virtual machines and Troubleshooting on VM Management.
  • Built performant, scalable ETL processes to load, cleanse and validate data
  • Participated in the full software development lifecycle with requirements, solution design, development, QA implementation, and product support using Scrum and other Agile methodologies
  • Expertise in Power BI, Power BI pro, Power BI Mobile.
  • Collaborate with team members and stakeholders in design and development of data environment
  • Preparing associated documentation for specifications, requirements, and testing

Environment: AWS, Gcp, Bigquery, Gcs Bucket, Power BI, Power BI pro, Power BI Mobile, G-Cloud Function, Apache Beam, Cloud Dataflow, Cloud Shell, Gsutil, Bq Command Line Utilities, Dataproc, Cloud Sql, Mysql, Data bricks, Posgres, Sql Server, Python, Scala, Spark, Hive, Spark-Sql

Data Engineer

Confidential, Negaunee, MI

Responsibilities:

  • Designed stream processing job used by Spark Streaming which is coded in Scala.
  • Ingested information from several sources like Kafka, Flume, and TCP sockets.
  • Processed data using advanced algorithms expressed with high-level functions like MapReduce, join and window.
  • Set up VirtualBox to gain access to a Linux environment. Also set up Vagrant which is crucial for setting up and installing the required software for running the Spark job.
  • The package job was first extracted to a deployment folder and then deployed to yarn so that yarn can take care of the scheduling and resource management.
  • Fed inbound events into the Scala-project-inbound topic in order to check if the window summary event functions as intended or not.
  • Gathered business requirements and converted them into new T-SQL stored procedures in visual studio for database project.
  • Performed unit tests on all code and packages.
  • Big data design and development using Apache Spark with Data Bricks and Azure Data Lake
  • Analyzed requirement and impact by participating in Joint Application Development sessions with business client online.
  • Performed and automated SQL Server version upgrades, patch installs and maintained relational databases.
  • Performed front line code reviews for other development teams.
  • Well versed experienced in Creating pipelines in Azure data factory using different activities like move and transform,Copy filter,foreach, Data Bricks etc.
  • Modified and maintained SQL Server stored procedures, views, ad-hoc queries, and SSIS packages used in the search engine optimization process.
  • Updated existing and created new reports using Microsoft SQL Server Reporting Services. Team consisted of 2 developers.
  • Created files, views, tables and data sets to support Sales Operations and Analytics teams
  • Monitored and tuned database resources and activities for SQL Server databases.

Environment: Visual Studio 2013 Dev12, SSMS .0.x.x, Power BI Desktop 2.19, Scala 2.12.3, Spark Streaming, Apache Hadoop 2.7.2, HDFS, YARN, slf4j 1.7.7

Data Analyst

Confidential

Responsibilities:

  • Developed stored procedures in MS SQL to fetch the data from different servers using FTP and processed these files to update the tables.
  • Responsible for Designing Logical and Physical data modeling for various data sources on Confidential Redshift.
  • Performed logical data modeling, physical Data modeling (including reverse engineering) using the Erwin Data modeling tool.
  • Created dimensional model for the reporting system by identifying required dimensions and facts using Erwin.
  • Designed and Developed ETL jobs to extract data from Salesforce replica and load it in data mart in Redshift.
  • Involved in performance tuning, stored procedures, views, triggers, cursors, pivot, unpivot functions, CTE's
  • Developed and delivered dynamic reporting solutions using SSRS.
  • Extensively used Erwin for Data modeling. Created Staging and Target Models for the Enterprise Data Warehouse.
  • Involved in Normalization / De normalization techniques for optimum performance in relational and dimensional database environments.
  • Resolved the data type inconsistencies between the source systems and the target system using the Mapping Documents and analyzing the database using SQL queries.
  • Worked on ETL testing, and used SSIS tester automated tool for unit and integration testing.
  • Designed and created SSIS/ETL framework from ground up.
  • Created new Tables, Sequences, Views, Procedure, Cursors and Triggers for database development.
  • Created ETL Pipeline using Spark and Hive for ingest data from multiple sources.
  • Involved in using SAP and transactions done in SAP - SD Module for handling customers of the client and generating the sales reports.
  • Creating reports using SQL Reporting Services (SSRS) for customized and ad-hoc Queries
  • Coordinated with clients directly to get data from different databases.
  • Worked on MS SQL Server, including SSRS, SSIS, and T-SQL.
  • Designed and developed schema data models.
  • Documented business workflows for stakeholder review.

Environment: ER Studio, SQL Server 2008, SSIS, Oracle, Business Objects XI, Rational Rose, Data stage, MS Visio, SQL, Crystal Reports 9

Java/J2EE Developer

Confidential

Responsibilities:

  • Responsible for understanding the scope of the project and requirement gathering.
  • Developed the web tier using JSP, Struts MVC to show account details and summary.
  • Created and maintained the configuration of the Spring Application Framework.
  • Implemented various design patterns - Singleton, Business Delegate, Value Object and Spring DAO.
  • Used Spring JDBC to write some DAO classes to interact with the database to access account information.
  • Mapped business objects to database using Hibernate.
  • Involved in writing Spring Configuration XML files that contains declarations and other dependent objects declaration.
  • Used Tomcat web server for development purpose.
  • Involved in creation of Test Cases for JUnit Testing.
  • Used Oracle as Database and used Toad for queries execution and also involved in writing SQL scripts, PL/ SQL code for procedures and functions.
  • Used CVS, Perforce as configuration management tool for code versioning and release.
  • Developed application using Eclipse and used build and deploy tool as Maven.
  • Used Log4J to print the logging, debugging, warning, info on the server console.

Environment: Java, J2EE, JSON, LINUX, XML, XSL, CSS, Java Script, PUTTY, Eclipse

We'd love your feedback!