Sr. Data Engineer Resume
SUMMARY:
- Overall, 8 years of technical IT experience in data analysis, design, development, and Implementation as a Data Engineer.
- Strong Experience in implementing Data warehouse solutions in Confidential Redshift; Worked on various projects to migrate data from on premise databases to Confidential Redshift, RDS and S3. Experience in Big Data analytics, Data manipulation, using Hadoop Eco system tools Map - Reduce, HDFS, Yarn/MRv2, Pig, Hive, HDFS, HBase, Spark, Kafka, Flume, Sqoop, Flume, Oozie, Avro, Sqoop, AWS, Spring Boot, Spark integration with Cassandra, Avro, Solr and Zookeeper. Proficiency in multiple databases like MongoDB, Cassandra, MySQL, ORACLE, and MS SQL Server. Worked on different file formats like delimited files, avro, Json and parquet. Docker container orchestration using ECS, ALB and lambda. Created Snowflake Schemas by normalizing the dimension tables as appropriate and creating a Sub Dimension named Demographic as a subset to the Customer Dimension. Experience in developing data pipelines using AWS services including EC2, S3, Redshift, Glue, Lambda functions, Step functions, Cloud Watch, SNS, Dynamo DB, SQS. Hands on experience in test driven development (TDD), Behavior driven development (BDD) and acceptance test driven development (ATDD) approaches. Managing
- Database, Azure Data Platform services (Azure Data Lake (ADLS), Data Factory(ADF), Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQL DB), SQLServer, Oracle, Data Warehouse etc. Build multiple Data Lakes. Extensive experience in Text Analytics, generating data visualizations using R, Python and creating dashboards using tools like Tableau, Power BI. Expertise in Java programming and have a good understanding on OOPs, I/O, Collections, Exceptions Handling, Lambda Expressions, Annotations Provided full life cycle support to logical/physical database design, schema management and deployment. Adept at database deployment phase with strict
EXPERIENCE:
Confidential
Sr. Data Engineer
Responsibilities:
- Worked with Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark - SQL, Data Frames and Pair RDD's. Performed advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark using Scala.*
- Developed Spark code using Scala and Spark-SQL for faster processing and testing. Created Spark jobs to do lighting speed analytics over the spark cluster. Extract Real time feed using Kafka and Spark Streaming and convert it to RDD and process data in the form of Data Frame and save the data as Parquet format in HDFS.* Responsible for writing real-time processing and core jobs using Spark Streaming with Kafka as a data pipe-line system and Configured Spark streaming to get ongoing information from the Kafka and store the stream information to HDFS.* Used Spark and Spark-SQL to read the parquet data and create the tables in hive using the Scala API. Involved in using the Spark application master to monitor the Spark jobs and capture the logs for the spark jobs. Developed multiple Kafka Producers and Consumers as per the software requirement specifications.* Used Spark Streaming APIs to perform transformations and actions on the fly for building common learner data model which gets the data from Kafka in near real time and persist it to Cassandra and built Real-time Data Pipelines with Kafka Connect and Spark Streaming.* Used Kafka and Kafka brokers, initiated the spark context and processed live streaming information with RDD and Used Kafka to load data into HDFS and NoSQL databases. Staging the Kafka Data into Snowflake DB by flattening the same for different functional service.* Developed stored procedures/views in Snowflake and Have Extracted and loaded data from AWS S3 to
- Snowflake Cloud Data Warehouse. Used Kafka functionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds and Created applications using Kafka, which monitors consumer lag within Apache Kafka clusters.* Developed end to end data processing pipelines that begin with receiving data using distributed messaging systems Kafka for persisting data into Cassandra. Responsible in development of Spark Cassandra connector to load data from flat file to Cassandra for analysis, modified Cassandra.yaml and Cassandra-env.sh files to set various configuration properties.* Used Sqoop to import the data on to Cassandra tables from different relational databases like Oracle, MySQL and Designed Column families in Cassandra performed data transformations, and then export the transformed data to Cassandra as per the business requirement.* Automated all the jobs starting from pulling the Data from different Data Sources like MySOL and pushing the result dataset to Hadoop Distributed File System and running MapReduce jobs and PIG/Hive using Oozie (Workflow management). Developed efficient MapReduce programs for fi
Confidential
Sr. Data Engineer
Responsibilities:
- Implemented Apache Airflow for authoring, scheduling, and monitoring Data Pipelines and designed several DAGs (Directed Acyclic Graph) for automating ETL pipelines. Performed data extraction, transformation, loading, and integration in data warehouse, operational data stores and master data management.* Have strong understanding of AWS components such as EC2 and S3 and responsible for data services and data movement infrastructures. Possesses knowledge in ETL concepts, building ETL solutions and Data modeling and have worked on architecting the ETL transformation layers and writing spark jobs to do the processing.* Worked on AWS Data Pipeline to configure data loads from S3 to into Redshift and used JSON schema to define table and column mapping from S3 data to Redshift. Wrote indexing and data distribution strategies optimized for sub - second query response* Designed & build infrastructure for the Google Cloud environment from scratch. Skillful in fact dimensional modeling (Star schema, Snowflake schema), transactional modeling and SCD (Slowly changing dimension). Leveraged cloud and GPU computing technologies for automated machine learning and analytics pipelines, such as AWS, GCP* Worked on confluence and Jira. Designed, and implemented configurable data delivery pipeline for scheduled updates to customer facing data stores built with Python* Proficient in Machine Learning techniques (Decision Trees, Linear/Logistic Regressors) and Statistical Modeling.
- Compiled data from various sources to perform complex analysis for actionable results. Knowledgeable in working with different join patterns and implemented both Map and Reduce Side Joins.* Wrote Flume configuration files for importing streaming log data into HBase with Flume. Imported several transactional logs from web servers with Flume to ingest the data into HDFS. Using Flume and Spool directory for loading the data from local system (LFS) to HDFS. Have Installed and configured pig, written Pig Latin scripts to convert the data from Text file to Avro format.* Created Partitioned Hive tables and worked on them using Hive QL and Worked on continuous Integration tools Jenkins and automated jar files at end of day. Worked with Tableau and Integrated Hive, Tableau Desktop reports and published to Tableau Server and Developed MapReduce programs in Java for parsing the raw data and populating staging Tables.* Skilled in setting up the whole app stack, setup, and debug log stash to send Apache logs to AWS Elastic search. Implemented Spark Scripts using Scala, Spark SQL to access hive tables into Spark for faster processing of data.* Extract Transform and Load data from Sources Systems to
- Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL and U-SQL Azure Data Lake Analytics. Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in Azure Databricks.
Confidential
Big Data Engineer
Responsibilities:
- Responsibilities include gathering business requirements, developing strategy for data cleansing and data migration, writing functional and technical specifications, creating source to target mapping, designing data profiling and data validation jobs in Informatica, and creating ETL jobs in Informatica.* Worked on Hadoop cluster which ranged from 4 - 8 nodes during pre-production stage and it was sometimes extended up to 24 nodes during production. And built APIs that will allow customer service representatives to access the data and answer queries. Designed changes to transform current Hadoop jobs to
- HBase.* Have Handled fixing of defects efficiently and worked with the QA and BA team for clarifications. Responsible for Cluster maintenance, Monitoring, commissioning and decommissioning Data nodes, Troubleshooting, Manage and review data backups, Manage & review log files. Extending the functionality of Hive with custom UDF s and UDAF's.* The new Business Data Warehouse (BDW) improved query/report performance reduced the time needed to develop reports and established self-service reporting model in Cognos for business users.* Implemented Bucketing and Partitioning using hive to assist the users with data analysis. Used Oozie scripts for deployment of the application and perforce as the secure versioning software. Executed Partitioning, Dynamic Partitions, Buckets in HIVE and Developed database management systems for easy access, storage, and retrieval of data.* Performed DB activities such as indexing, performance tuning, and backup and restore. Expertise in writing Hadoop Jobs for analysing data using Hive QL (Queries), Pig Latin (Data flow language), and custom MapReduce programs in Java.* Have executed various performance optimizations like using distributed cache for small
- Partition, Bucketing in the hive and Map Side joins. Expert in creating Hive UDFs using Java to analyse the data efficiently.* Responsible for loading the data from BDW Oracle database, Teradata into HDFS using Sqoop. Implemented AJAX, JSON, and Java script to create interactive web screens. Wrote data ingestion systems to pull data from traditional RDBMS platforms such as Oracle and Teradata and store it in NoSQL databases such as MongoDB. Involved in loading and transforming large sets of Structured, Semi-Structured and Unstructured data and analysed them by running Hive queries. Processed the image data through the Hadoop distributed system by using Map and Reduce then stored into HDFS.* Created Session Beans and controller Servlets for handling HTTP requests from Talend. Performed Data Visualization and Designed Dashboards with Tableau and generated complex reports including chars, summaries, and graphs to interpret the findings to the team and stakeholders.* Wrote documentation for each report including purpose, data source, column mapping, transformation, and user group. Utilized Waterfall methodology for team and
Confidential
Data Engineer
Responsibilities:
- Skilled in Big Data Analytics and design in Hadoop ecosystem using MapReduce Programming, Spark, Hive, Pig, Sqoop, HBase, Oozie, Impala, Kafka. Build the Oozie pipeline which performs several actions like file move process, Sqoop the data from the source Teradata or SQL and exports into the hive staging tables and performing aggregations as per business requirements and loading into the main tables.* Running of Apache Hadoop, CDH and Map - R distros, dubbed Elastic MapReduce (EMR) on (EC2). Performing the forking action whenever there is a scope of parallel process for optimization of data latency.* Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python. Performed pig script which picks the data from one Hdfs path and performs aggregation and loads into another path which later pulls populates into another domain table. Converted this script into a jar and passed as parameter in Oozie script* Developed JSON Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the SQL Activity. Build an ETL which utilizes spark jar inside which executes the business analytical model.* Professional on git bash commands like git pull to pull the code from source and developing it as per the requirements, git add to add files, git commit after the code build and git push to the pre prod environment for the code review and later used screwdriver. yaml which actually build the code, generates artifacts which releases into production*
- Created logical data model from the conceptual model and its conversion into the physical database design using Erwin. Involved in transforming data from legacy tables to HDFS, and HBase tables using Sqoop. Connected to AWS Redshift through Tableau to extract live data for real time analysis.* Developed Data mapping, Transformation and Cleansing rules for the Data Management involving OLTP and OLAP. Involved in creating UNIX shell Scripting. Defragmentation of tables, partitioning, compressing and indexes for improved performance and efficiency.* Developed reusable objects like PL/SQL program units and libraries, database procedures and functions, database triggers to be used by the team and satisfying the business rules. Used SQL Server Integrations Services (SSIS) for extraction, transformation, and loading data into target system from multiple sources* Developed and implemented R and Shiny application which showcases machine learning for business forecasting. Developed predictive models using Python & R to predict customers churn and classification of customers. Partner with infrastructure and platform teams to configure, tune tools, automate tasks and guide the evolution of internal big data ecosystem; serve as a bridge between data scientists and infrastructure/platform teams.* Implemented Big Data Analytics and Advanced Data Science techniques to identify trends, patterns, and discrepancies on petabytes of
Confidential
Data Analyst
Responsibilities:
- Imported Legacy data from SQL Server and Teradata into Amazon S3. Created consumption views on top of metrics to reduce the running time for complex queries. Exported Data into Snowflake by creating Staging Tables to load Data of different files from Amazon S3.* Compare the data in a leaf level process from various databases when data transformation or data loading takes place. I need to analyze and look into the data quality when these types of loads are done (To look for any data loss, data corruption).* As a part of Data Migration, wrote many SQL Scripts for Mismatch of data and worked on loading the history data from Teradata SQL to snowflake. Developed SQL scripts to Upload, Retrieve, Manipulate and handle sensitive data (National Provider Identifier Data I.e. Name, Address, SSN, Phone No) in Teradata, SQL Server Management Studio and Snowflake Databases for the Project* Worked on to retrieve the data from FS to S3 using spark commands. Built S3 buckets and managed policies for S3 buckets and used S3 bucket and Glacier for storage and backup on AWS. Created performance dashboards in Tableau/ Excel / Power point for the key stakeholders* Incorporated predictive modelling (rule engine) to evaluate the
Customer/Seller health score using python scripts, performed computations, and integrated with the Tableau viz. Worked with stakeholders to communicate campaign results, strategy, issues or needs.* Analysed marketing campaigns from various perspectives including CTR, conversion rates, seasonal/geographical trends, search queries, landing page, conversion funnel, quality score, competitors, distribution channel, etc. to achieve maximum ROI for clients.* Worked with business to identify the gaps in mobile tracking and come up with the solution to solve. Analysed click events of Hybrid landing page which includes bounce rate, conversion rate, Jump back rate, List/Gallery view, etc. and provide valuable information for landing page optimization.* Evaluated the traffic and performance of Daily deals PLA ads and compare those items with non - daily deal items to see the possibility of increasing ROI. suggested improvements and modify existing BI components (Reports, Stored Procedures). Understood Business requirements to the core and Came up with Test Strategy based on Business rules. Prepared Test Plan to ensure QA and Development phases are in parallel* Written and executed Test Cases and reviewed with Business & Development Teams. Implemented Defect Tracking process using JIRA tool by assigning bugs to Development Team. Performed Automated Regression tool (Qute) and reduced manual effort and increased team productivity* Involved in Functional Testing, Integration testing, Regression Testing, Smoke testing and performance Testing. Tested Hadoop MapReduce developed in python, pig, Hive. Created Metric tables, End user views in Snowflake to feed data for Tableau refresh.* Generated Custom SQL to verify th
