Data Engineer Resume
Upper Gwynedd, PA
SUMMARY
- Reliable & experienced Data and Business Intelligence Engineer, with 8+ years of experience in Cloud/On Premise presents distinctive competency in Business Analytics/ Business Intelligence/ Data Warehouse /Data Marts/Big Data/ Data Analytics/ Customer Relationship Management (CRM)/ Marketing Relationship Management (MRM)/ Sales & Marketing/ Supply Chain Management, Skilled in Python, AWS Cloud, AZURE Data Lake, GCP, Data Model, Data Warehouse and BI Reporting with strong emphasis on Data Engineering and Data Analysis using Big Data tools and application development using HADOOP framework and related technologies such as HDFS, MapReduce, HIVE, PIG, HBASE, STORM, YARN, OOZIE, SQOOP, Air Flow and Zookeeper and includes working experience in Spark Core, Spark SQL, Spark Streaming, Scala, Python and Kafka.
- Experience gathering customer requirements, writing test cases, and partnering with developers to ensure full understanding of internal/external customer needs.
- Excellent knowledge in building data engineering pipelines, automating, and fine - tuning for both batch and real time data pipelines.
- Expertise in writing end to end Data processing Jobs to analyze data using MapReduce, Spark, and Hive.
- Experience with Apache Spark ecosystem using Spark-Core, SQL, Data Frames, RDD's, and knowledge on Spark MLLib.
- Experienced in data manipulation using python for loading and extraction as well as with python libraries such as NumPy, SciPy and Pandas for data analysis and numerical computations.
- Experience on Migrating SQL database to Azure data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse,GCP and controlling and granting database access and Migrating On premise databases to Azure Data Lake store using Azure Data factory.
- Hands-on experience withAmazon EC2, S3, RDS, IAM, Auto Scaling, CloudWatch, SNS, Athena, Glue, Kinesis, Lambda, EMR, Redshift, DynamoDB and other services of the AWS family.
- Built real timedata pipelinesby developingKafkaproducers andstreaming applications for consuming.
- Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, experienced in maintaining the Hadoop cluster on AWS EMR.
- Designed and developed logical and physical data models that utilize concept such as Star Schema, Snowflake Schema and Slowly Changing Dimensions.
- Design and develop Spark applications using Pyspark and Spark-SQL for data extraction, transformation, and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage pattern.
- Experience in Work on AWS Databases like Elastic Cache (Memcached & Redis) and NoSQL databases HBase, Cassandra & MongoDB for database performance tuning & data modeling.
- Experience in Azure Cloud Services (PaaS & IaaS), Azure Synapse Analytics, SQL Azure, Data Factory, Azure Analysis services, Application Insights, Azure Monitoring, Key Vault, Azure Data Lake.
- Strong experience in using Spark Streaming, Spark SQL, and other components of spark like accumulators, Broadcast variables, different levels of caching and optimization techniques for spark jobs
- Good experience working on AWS-Bigdata/Hadoop Ecosystem in the implementation of Data Lake.
- Extensive experience and hands-on knowledge across the catalog of GCP technologies, GCP Cloud, Big Query, SQL, Data-flows, Databases Oracle, DB2, SQL Server, Jenkins - CI/CD, Java.
- Expertise in designing/ developing dynamically scalable, highly available, highly reliable and fault-tolerant applications on GCP.
- Demonstrated development & implementation experience in building scalable Cloud Applications on GCP.
- Strong Hadoop and platform support experience with all the entire suite of tools and services in majorHadoop Distributions- Cloudera, Amazon EMR, Azure HDInsight, and Hortonworks.
- Experience in writingREST APIsinPythonfor large-scale applications.
- Hands-on experience in setting up workflow using Apache Airflow and Oozie workflow engine for managing and scheduling Hadoop jobs.
- Experience in data warehousing and business intelligence area in various domain.
- Created Tableau dashboards designing with large data volumes from data source SQL servers.
- Extract, Transform and Load (ETL) source data into respective target tables to build Data Marts.
- Active involvement in all scrum ceremonies - Sprint Planning, Daily Scrum, Sprint Review and Retrospective meetings and assisted Product owner in creating and prioritizing user stories.
TECHNICAL SKILLS
Hadoop/Spark Ecosystem: Hadoop, MapReduce, Pig, Hive/impala, YARN, Kafka, Flume, Oozie, Zookeeper, Spark, Airflow
Hadoop Distribution: Cloudera distribution and Hortonworks
Cloud Platforms: AWS: Amazon EC2, S3, RDS, IAM, Auto Scaling, CloudWatch, SNS, Athena, Glue, Kinesis, Lambda, EMR, Redshift, DynamoDB Azure: Azure Cloud Services (PaaS & IaaS), Azure Synapse Analytics, SQL Azure, Data Factory, Azure Analysis services, Application Insights, Azure Monitoring, Key Vault, Azure Data Lake, Azure HDInsight GCP, OpenStack.
ETL/BI Tools: Informatica, SSIS, Tableau, PowerBI, SSRS
CI/CD: Jenkins, Splunk, Ant, Maven, Gradle.
Ticketing Tools: JIRA, Service Now, Remedy
Operating Systems: Linux, Windows, Ubuntu, Unix
Database: Oracle, SQL Server, Cassandra, Teradata, PostgreSQL, Snowflake, HBase, MongoDB
Programming Languages: Scala, Hibernate, PL/SQL, R
Scripting: Python, Shell Scripting, JavaScript, jQuery, HTML, JSON, XML.
Web/Application server: Apache Tomcat, WebLogic, WebSphere Tools Eclipse, NetBeans
BI Tools: Tableau Desktop 10.x/9.x/ 8.3/7, Tableau Server 8.2, OBIEE 11g/10.1.3.4, PowerBI2.73
Machine Learning And Statistics: Regression, Random Forest, Clustering, Time-Series Forecasting, HypothesisExplanatory Data Analysis
Version Control: Git, Subversion, Bitbucket, TFS.
SDLC: Agile, Scrum, Waterfall, Kanban.
PROFESSIONAL EXPERIENCE
Confidential, Upper Gwynedd, PA
Data Engineer
Responsibilities:
- Working in Azure Development on Azure web application, App services, Azure storage, Azure SQL Database, Azure Virtual Machines, Azure AD, Azure search, Azure DNS and Azure VPN Gateway.
- Developed multiple applications required for transforming data across multiple layers of Enterprise Analytics Platform and implement Big Data solutions to support distributed processing using Big Data technologies.
- Responsible for data identification and extraction using third-party ETL and data-transformation tools or scripts. (e.g., SQL, Python)
- Worked on migration of data from On-prem SQL server to Cloud databases (Azure Synapse Analytics (DW) & Azure SQL DB).
- Developed and managed Azure Data Factory pipelines that extracted data from various data sources, transformed it according to business rules, using python scripts that utilized Pyspark and consumed APIs to move data into an Azure SQL database.
- Created a new data quality check framework project in Python that utilized pandas.
- Implemented source control for Azure Data Factory pipelines utilizing Azure Repos.
- Created Hive/Spark external tables for each source table in the Data Lake and written Hive SQL and Spark SQL to parse the logs and structure them in tabular format to facilitate effective querying log data.
- Designed and developed ETL & ETL frameworks using Azure Data Factory and Azure Data Bricks.
- Created generic data bricks NOTEBOOKs for performing data cleansing.
- Created Azure Data factory pipelines to refactor on-prem SSIS packages into Data factory pipelines.
- Working with Azure BLOB and Data Lake storage for loading data into Azure SQL Synapse (DW).
- Ingested and transformed source data using Azure Data flows and Azure HDInsight.
- Created Azure Functions to ingest data at regular intervals.
- Created Data Bricks notebooks for performing complex transformations and integrated them as activities in ADF pipelines.
- Written complex SQL queries for data analysis and extraction of data in required format.
- Created Power BI DataMart’s and reports for various stakeholders in the business.
- Created CI/CD pipelines using Azure DevOps.
- Enhanced the functionality of existing ADF pipeline by adding new logic to transform the data.
- Worked on Spark jobs for data preprocessing, validation, normalization, and transmission.
- Optimized code and configurations for performance tuning of Spark jobs.
- Worked with unstructured and semi structured data sets to aggregate and build analytics on the data.
- Work independently with business stakeholders with strong emphasis on influencing and collaboration.
- Daily participation in Agile based Scrum team with tight deadlines.
Environment: Azure Synapse Analytics, Azure Data Factory, Azure Data bricks, Azure Synapse Studio, Hadoop, SQL Server, Power BI, Oracle 12c/11g, SQL scripting, PL/SQL, Python, Unix Shell, Jira, Confluence.
Confidential, Detroit, MI
Data Engineer/Data Analyst
Responsibilities:
- Created Entity Relationship Diagrams (ERD), Functional diagrams, Data flow diagrams and enforced referential integrity constraints and created logical and physical models using Erwin.
- Analyzed the system for new enhancements/functionalities and perform Impact analysis of the application for implementing ETL changes.
- Worked on AWS EMR to transform and move large amounts ofdatainto and out of otherAWSdatastores and databases, Amazon Simple Storage Service (Amazon S3) and DynamoDB.
- Migrated on premise database structure to Confidential Redshift data warehouse
- Defined and deployed monitoring, metrics, and logging systems on AWS.
- Connected to Amazon Redshift through Tableau to extract live data for real time analysis.
- Implemented lambda functions to extract the data from an API and load them into Dynamo DB.
- Configured step functions to orchestrate the multiple EMR tasks for data processing.
- UsedSpark-SQLto LoadJSONdata and createSchema RDDand loaded it intoHiveTables and handled Structured data usingSparkSQL.
- Imported data fromAWS S3and intoSparkRDDand performed transformations and actions onRDD's.
- Converted Hive/SQL queries into Spark Transformations using Spark RDDs and Scala and involved in using SQOOP for importing and exporting data between RDBMS and HDFS.
- Optimizing and tuning the Redshift environment, enabling queries to perform up to 100x faster for Tableau and SAS Visual Analytics.
- Designed solutions to process high volume data stream ingestion, processin low latency data provisioning using Hadoop Ecosystems Hive, Pig, Scoop, Kafka, Python, Spark, Scala, NoSql, Nifi, and Druid.
- Designed and implemented big data ingestion pipelines to ingest multi-TB data from various data source using Kafka, Spark streaming including data quality checks, transformation, and stored as efficient storage formats Performing data wrangling for a variety of downstream purposes such as analytics using PySpark.
- Implemented Workload Management (WML) in Redshift to prioritize basic dashboard queries over more complex longer running adhoc queries. This allowed for a more reliable and faster reporting interface, giving sub-second query response for basic queries.
- Implemented a Continuous Delivery pipeline with Docker, GitHub, and AWS.
- Worked on Development & implementation experience in building scalable Cloud Applications on GCP.
- Created ad hoc queries and reports to support business decisions SQL Server Reporting Services (SSRS).
- Analyze the existing application programs and tune SQL queries using execution plan, query analyzer, SQL Profiler, and database engine tuning advisor to enhance performance.
- Scheduled Airflow DAGs to run multiple Hive and Pig jobs, which independently run with time and data availability and Performed Exploratory Data Analysis and Data Visualizations using Python, and Tableau.
Environment: Hadoop/Bigdata Ecosystem (Spark, Kafka, Hive, HDFS, Sqoop, Oozie, Cassandra, MongoDB), AWS (S3, AWS Glue, GCP Redshift, RDS, Lambda, Athena, SNS, SQS), Oracle, Docker, Git, SQL Server, Python 3.x, Pyspark, Teradata, Tableau, Quick sight, Data warehousing.
Confidential, Los Angeles, CA
Big Data Developer/ Data Analyst
Responsibilities:
- Worked on Hadoop eco-systems including Hive, HBase, Oozie, Pig, Zookeeper, Spark Streaming MCS (MapR Control System) and so on with MapR distribution.
- Installed and configured Hadoop MapReduce, HDFS, Developed multiple MapReduce jobs in Java for data cleaning and pre-processing.
- Built code for real time data ingestion using Java, MapR-Streams (Kafka) and STORM.
- Involved in development of Hadoop System and improving multi-node Hadoop Cluster performance.
- Worked on analyzing Hadoop stack different big data tools including Pig, Hive, HBase database Sqoop.
- Developed data pipeline using flume, Sqoop and pig to extract the data from weblogs and store in HDFS
- Worked with different data sources like Avro data files, XML files, JSON files, SQL server and Oracle to load data into Hive tables.
- Used Spark to create the structured data from large amount of unstructured data from various sources.
- Implemented usage of Amazon EMR for processing Big Data across Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3).
- Performed transformations, cleaning and filtering on imported data using Hive, MapReduce, Impala and loaded final data into HDFS.
- Developed Python scripts to find vulnerabilities with SQL Queries by doing SQL injection.
- Experienced in designing and developing POC’s in Spark using Scala to compare the performance of Spark with Hive and SQL/Oracle.
- Specified the cluster size, allocating Resource pool, Distribution of Hadoop by writing the specification texts in JSON File format.
- Imported weblogs & unstructured data using the Apache Flume and stores the data in Flume channel.
- Exported event weblogs to HDFS by creating a HDFS sink which directly deposits the weblogs in HDFS.
- Used RESTful web services with MVC for parsing and processing XML data.
- Collaborated and communicated the results of analysis to the decision makers by presenting actionable insights by using visualization charts and dashboards in Amazon Quick Sight.
Environment: Hadoop, Apache Spark, HDFS, Hive, Spark SQL, Pyspark, Python, Django, Oracle SQL, Tableau, AWS, Hadoop distribution of Horton Works, Cloudera, Pig, HBase, Linux, XML, Zookeeper
Confidential, Denver, CO
BI Developer / Data Analyst
Responsibilities:
- Reinforced the development and implementing of databases (ETL) in SSIS of a pharmaceutical manufacturing company to optimize the efficiency of the system and operations by 8% along with the identification of the KPIs
- Analyzed historical data by data cleaning and exploration in Python and achieved 89% accuracy by converting data into actionable insights, interpreted results using statistical techniques along with predicting future outcomes
- Built reports using SQL Server Reporting (SSRS), developed interactive dashboards in Excel and Tableau to visualize KPIs and provided business recommendations that increased the revenue by 15%
- As a Data Visualization Consultant supporting Risk Consulting Team in Model Documentation and Data Visualizations for the Data Science Team which develops Credit and Market Risk Predictive Models for the Banking clients.
- Worked as a Tableau Desktop Developer focusing on developing high-end visualizations driven by data coming in from various data sources including Flat files, SQL Server, and MS Excel.
- Developed various stories and dashboards using multiple data sources and by blending them on a single worksheet in Tableau Desktop version.
- Involved extensively in dashboards development such as creating Tableau Extracts, Tableau Connectors (Live and Extract), formatting and report operations (sorting, filtering, ranking, Top-N Analysis).
- Worked extensively with Advance analysis Actions, Calculations, Parameters, Background images, Maps, Trend Lines, Statistics, and Log Axes. Groups, hierarchies in Tableau.
- Designed, Deployed, and Integrated and Maintained MDM systems.
- Used advanced data visualization and representation techniques in Tableau to provide an easy-to-understand interface for end users to quickly identify key areas within their data.
- Trained end users in various parts of the world to effectively use Tableau.
- Create and enhance the standard of client communications like company reports & screen presentations in PowerPoint, various Reports in MS Excel/Tableau through corporate template as per specification.
- Additionally, performing the role of the Quality Audit on presentation for the new trainees.
- Trained end users in various parts of the world to effectively use tableau and customize the reports based on the department and needs of the users.
- Engaged in the design and production of visual communication materials like charts, financial presentations based on PowerPoint/Think cell/Tableau for use of Confidential Leadership Team.
Environment: Tableau, MS SQL Server, MS Excel, Flat files, MS PowerPoint, SSRS, SSIS, Azure
Confidential
Data Analyst
Responsibilities:
- Actively participated in logical design of database design to meet new product requirement using ERWIN.
- Extensively used SQL queries to check storage and accuracy of data in database tables and utilized SQL for querying the SQL database.
- Designed easy to follow visualizations using Tableau software and published dashboards on web and desktop platforms.
- Worked with Business Owners on perfecting the process, increasing the efficiency of the systems.
- Conducted analysis, developed information systems, and generated accurate and comprehensive reports on operational data.
- As a Data Analyst, worked closely with Business Analysts to gather requirements and design a reliable and scalable data pipelines
- Power BI dashboard maintenance, SQL Tableau Python Data Handling
- Used Microsoft Excel (e.g., charts, filters, vlookup, pivot tables)
- Created and published multiple dashboards and reports using Tableau server.
- Worked on both batch processing and streaming data Sources. Used Spark streaming and Kafka for the streaming data processing.
- Preparing Dashboards using calculations, parameters in Tableau
- Working knowledge on table design anddata management using HDFS, Hive, Impala, Sqoop, MySQL, and Kafka.
- Developed python scripts for data cleaning, analysis and automating day to day activities.
- Developed automated reports using Tableau, Python and MySQL to reduce the manual intervention saving 20 hours a month
- Collected, calculated, and checked information from various data sources to devise solutions, design systems, and implement information management processes.
- Cooperated and communicated well with other personnel across departments to explain and assist in the integration of information management and data communication systems.
- Recommended other methods and technology derived from the gathered data to maximize the efficiency of project implementation.
Environment: SQL Server, SSIS, SSRS, Tableau, Erwin, Flat Files, Power BI, Microsoft Dynamics CRM, SQL Server 2016, SQL Server Management Studio (SSMS), Visual Studio 2015, T-SQL, Excel, DAX.
