We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

0/5 (Submit Your Rating)

San Diego, CA

SUMMARY

  • 8 years of overall experience with strong emphasis on Design, Development, Implementation, Testing and Deployment of Software Applications, Over 4+ years of comprehensive IT experience in BigData Analytics, Hadoop, HDFS, MapReduce, YARN, Hadoop Ecosystem and Shell Scripting.
  • Proficient in SQL Server and T - SQL (DDL and DML) in constructing Tables, Normalization/ De normalization Techniques on database Tables.
  • Experience in Creating and Updating Clustered and Non-Clustered Indexes to keep up the SQL Server Performance in OLTP and OLAP environments.
  • Experience across entire Microsoft suite of products including SQL Server, ETL (SSIS), Report Design (SSRS), and Multidimensional cubes development (SSAS), MDS (Master Data Services), DQS (Data Quality services), Azure Data Lake, Azure Data Factory and Power BI SDLC from prototyping to deployment.
  • Highly capable for processing large sets of Structured, Semi-structured and Unstructured datasets and supporting Big Data applications.
  • Hands on experience in Azure Development, worked on Azure web application, App services, Azure storage, Azure SQL Database, Virtual machines, Fabric controller, Azure AD, Azure search, and notification hub.
  • Designed, configured, and deployed Microsoft Azure for a multitude of applications utilizing the Azure stack (Including Compute, Web & Mobile, Blobs, Resource Groups, Azure SQL, Cloud Services, and ARM), focusing on high - availability, fault tolerance, and auto-scaling.
  • Expertise in Microsoft Azure Cloud Services (PaaS & IaaS), Application Insights, Document DB, Internet of Things (IoT), Azure Monitoring, Key Vault, Visual Studio Online (VSO) and SQL Azure.
  • Experience on Migrating SQL database to Azure data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and Migrating On premise databases to Azure Data Lake store using Azure Data factory.
  • Extensive involvement in Designing Azure Resource Manager Template and in designing custom build steps using PowerShell.
  • Designed and developed Cloud Service projects and deployed to Web Apps, PaaS, and IaaS.
  • Configured SQL Server Master Data Services (MDS) in Windows Azure IaaS. Manage different AZURE environment for provisioning of Linux servers and services executed by the providers.
  • Good at Manage hosting plans for Azure Infrastructure, implementing & deploying workloads on Azure virtual machines (VMs).
  • Hands-on experience with Amazon EC2, Amazon S3, Amazon RDS, VPC, IAM, Amazon Elastic Load Balancing, Auto Scaling, CloudWatch, SNS, SES, SQS, Lambda, EMR and other services of the AWS family.
  • Experience in Apache Spark cluster and streams processing using Spark Streaming.
  • Expertise in moving large amounts of log, streaming event data and Transactional data using Flume.
  • Expertise in writing Pig Latin, Hive Scripts and extended their functionality using User Defined Functions (UDF's).
  • Experience in Developing Spark applications using Spark - SQL for data extraction, transformation, and aggregation from multiple file formats for analysing & transforming the data to uncover insights into the customer usage patterns.
  • Hands on experience in installing, configuring, and using Hadoop ecosystem components like HDFS, MapReduce, HBase, Zookeeper, Oozie, Hive, Sqoop, Pig, and Kafka.
  • Experience in Hadoop Shell commands, writing MapReduce Programs, verifying managing, and reviewing Hadoop Log files.
  • Expertise in transferring data between a Hadoop ecosystem and structured data storage in a RDBMS such as MY SQL, Oracle, Teradata and DB2 using Sqoop.
  • Experience in NoSQL databases like Mongo DB, HBase and Cassandra.
  • Experience in working with Transactional Databases like Oracle, SQL server, My SQL, and Db2.
  • Expertise in developing SQL queries, Stored Procedures, and excellent development experience with Agile Methodology and have ability to adapt to evolving technology, Strong sense of Responsibility and Accomplishment.

TECHNICAL SKILLS

Operating Systems: Windows, Linux

Big Data Eco System: Hadoop, HDFS, MapReduce, Hive, Pig, HBase, Spark, Scala, Impala, kafka, Hue, Sqoop, Oozie, Flume, Zookeeper, Cassandra, Cloudera CDH5, Azure, Azure Databricks, Azure SQL database, Azure SQL Datawarehouse.

Relational/NOSQL Databases: Oracle, MySQL, SQL Server, DB2, Mongo DB, Teradata, HBase, Cassandra.

Cloud Technologies: Azure, AWS

SDLC Methodologies: Agile/Scrum, Waterfall.

Version Control Tools: SVN, GitEnvironment: Hadoop, HDFS, Map Reduce, Oozie, MySQL, Cloudera, Hive, HBase, Teradata, UNIX Shell Scripting, Pig, Hive, Windows.

PROFESSIONAL EXPERIENCE

Confidential, San Diego, CA

Sr. Data Engineer

Responsibilities:

  • Responsible for Managing, Analysing and Transforming Petabytes of data and quick validation check on FTP file arrival from S3 Bucket to HDFS.
  • Responsible for analysing large data sets and derive customer usage patterns by developing new MapReduce programs.
  • Creation of Hive tables and loading data incrementally into the tables using Dynamic Partitioning and Worked on Avro Files, JSON Records.
  • Worked on Hive by creating external and internal tables, loading it with data and writing Hive queries.
  • Involved in development and usage of UDTF's and UDAF's for decoding Log Record Fields and Conversion's, Generating Minute Buckets for the specified Time Interval's and JSON Field Extractor.
  • Developed Pig and Hive UDF's to analyse the complex data to find specific user behaviour.
  • Developed workflow in Oozie to automate the tasks of loading data into HDFS and pre-processing with Pig and Hive.
  • Involved for Cassandra Database Schema design.
  • Using BULK LOAD Utility data pushed to Cassandra databases.
  • Responsible for Scheduling using Active Batch jobs and Cron jobs.
  • Actively updated the upper management with daily updates on the progress of project that include the classification levels that were achieved on the data.
  • Performed end- to-end Architecture & implementation assessment of various AWS services like Amazon EMR, Redshift, S3.
  • Used AWS EMR to transform and move large amounts of data into and out of other AWS data stores and databases, such as Amazon Simple Storage Service (Amazon S3) and Amazon DynamoDB.
  • Implemented AWS Step Functions to automate and orchestrate the Amazon SageMaker related tasks such as publishing data to S3, training ML model and deploying it for prediction.
  • Integrated Apache Airflow with AWS to monitor multi-stage ML workflows with the tasks running on Amazon SageMaker, AWS EMR, S3, RDS, Redshift, Lambda, Boto3, DynamoDB, Amazon SageMaker, Apache.
  • Worked with file formats TEXT, AVRO, PARQUET, and SEQUENCE files.

Confidential, Hornell, NY

Sr. Data Engineer

Responsibilities:

  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL and U-SQL Azure Data Lake Analytics. Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in In Azure Databricks.
  • Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards. analyze, design, and build Modern data solutions using Azure PaaS service to support visualization of data. Understand current Production state of application and determine the impact of new implementation on existing business processes.
  • Experience in managing Azure Storage Accounts. Wrote many stored procedures for cleaning, manipulating, and processing data between the databases.
  • Developed a process for Sqooping data from multiple sources like SQL Server, Oracle, and Teradata.
  • Responsible for creation of mapping document from source fields to destination fields mapping.
  • Developed a shell script to create staging, landing tables with the same schema as the source and generate the properties which are used by Oozie jobs.
  • Developed Oozie workflows for executing Sqoop and Hive actions.
  • Worked with NoSQL databases like HBase in creating HBase tables to load large sets of semi structured data coming from various sources.
  • Developed scripts to run Oozie workflows, capture the logs of all jobs that run on cluster and create a metadata table which specifies the execution times of each job.
  • Developed Hive scripts for performing transformation logic and loading the data from staging zone to final landing zone.
  • Involved in loading transactional data into HDFS using Flume for Fraud Analytics.
  • Import and Export of data using Sqoop between MySQL to HDFS on regular basis.
  • Responsible for developing multiple Kafka Producers and Consumers from scratch as per the software requirement specifications.
  • Involved in migrating the needed data from Oracle, MySQL in to HDFS using Sqoop and importing various formats of flat files in to HDFS.

Confidential, East Hanover, NJ

Data Engineer

Responsibilities:

  • Performed performance tuning and troubleshooting of MapReduce jobs by analysing and reviewing Hadoop log files.
  • Involved Low level design for MR, Hive, Impala, Shell scripts to process data.
  • Involved in complete Big Data flow of the application starting from data ingestion upstream to HDFS, processing the data in HDFS and analysing the data.
  • Creating Hive tables to import large data sets from various relational databases using Sqoop and export the analysed data back for visualization and report generation by the BI team.
  • Installing and configuring Hive, Sqoop, Flume, Oozie on the Hadoop clusters.
  • Involved in scheduling Oozie workflow engine to run multiple Hive and Pig jobs.
  • Implemented partitioning, dynamic partitions and buckets in HIVE.
  • Developed Hive Scripts to create the views and apply transformation logic in the Target Database.
  • Involved in the design of Data Mart and Data Lake to provide faster insight into the Data.
  • Involved in using Stream Sets Data Collector tool and created Data Flows for one of the streaming applications.
  • Used Kafka as a data pipeline between JMS (Producer) and Spark Streaming Application (Consumer).
  • Designed and Maintained Oozie workflows to manage the flow of jobs in the cluster.

Confidential

Analyst

Responsibilities:

  • Design and develop ETL data pipeline using Spark App to fetch data from Legacy system and third-party API, social media sites.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
  • Design and develop ETL data pipeline using Spark App to fetch data from Legacy system and third-party API, social media sites.
  • Worked on spark applications and launching clusters with spark in EMR console.
  • Importing the data into Spark from Kafka Consumer group using Spark Streaming APIs.
  • Implemented Spark RDD transformations to map business analysis and apply actions on top of transformations.
  • Used Microsoft Azure for building the applications and for building, testing, deploying the applications.
  • Worked with SQOOP import and export functionalities to handle large data set transfer between MySQL database and HDFS.
  • Develop spark SQL tables & queries to perform Adhoc data analytics for analyst team.
  • Monitoring Spark clusters.
  • Involved in developing and designing POCs using Scala and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata.

Confidential

Programmer Analyst

Responsibilities:

  • Working on complete life cycle of software development, which included new requirement gathering, redesigning, and implementing the business specific functionalities, testing, and assisted in deployment of the project to the PROD environment.
  • Load log data into HDFS using Flume. Worked extensively in creating MapReduce jobs to power data for search and aggregation.
  • Installed and Configured Cloudera Hadoop CDH4 via Cloudera Manager in a pseudo distributed mode and cluster mode as a proof of concept.
  • Involved in creating Hive tables, loading with data, and writing hive queries which will run internally in map reduce way.
  • Implemented Spark using Scala and utilizing Spark core, Spark streaming and Spark SQL API for faster processing of data instead of Map reduce in Java and Loaded data into the cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.

We'd love your feedback!