Sr. Data Engineer Resume
FloridA
SUMMARY
- Around 7+ years of experience as a Data Engineer and extensively worked with designing, developing, and implementing Data models for enterprise - level applications and BI solutions.
- Experience in designing and building Data Management Lifecycle covering Data Ingestion, Data integration, Data consumption, Data delivery, and integration Reporting, Analytics, and System-System integration.
- Proficient in Big Data environment and Hands-on experience in utilizing Hadoop environment components for large-scale data processing including structured and semi-structured data.
- Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
- Extensive experience with Azure cloud technologies like Azure Data Lake Storage, Azure Data Factory, Azure SQL, Azure Data Warehouse, Azure Synapse Analytical, Azure Analytical Services, Azure HDInsight, and Databricks.
- Solid Knowledge of AWS services like AWS EMR, Redshift, S3, EC2, and concepts, configuring the servers for auto-scaling and elastic load balancing.
- Experience with monitoring the web services using Hadoop and Spark for controlling the applications and analyzing their operation and performance.
- Experienced in Python data manipulation for loading and extraction as well as with Python libraries such as NumPy, Pandas, and SciPy for data analysis and numerical computations.
- Good knowledge and experience with NoSQL databases like HBase, Cassandra, and MongoDB and SQL databases like Teradata, Oracle, PostgreSQL, and SQL Server.
- Experience in the development and design of various scalable systems using Hadoop technologies in various environments and analyzing data using MapReduce, Hive, and PIG.
- Hands-on use of Spark and Scala to compare the performance of Spark with Hive and SQL, and Spark SQL to manipulate Data Frames in Scala.
- Strong knowledge in working with ETL methods for data extraction, transformation, and loading in corporate-wide ETL Solutions and Data Warehouse tools for reporting and data analysis.
- Hands-on experience in designing and implementing data engineering pipelines and analyzing data using Hadoop ecosystem tools like HDFS, Spark, Sqoop, Hive, Flume, Kafka, Impala, PySpark, Oozie, and HBase.
- Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
- Experience with different ETL tool environments like SSIS, Informatica, and reporting tool environments like SQL Server Reporting Services, and Business Objects.
- Experience in deployment of applications and scripting using the Unix/Linux Shell scripting.
- Solid knowledge of Data Marts, Operational Data Store, OLAP, Dimensional Data Modeling with Star Schema Modeling, Snow Flake Modeling for Dimensions Tables using Analysis Services.
- Extensive experience with various databases like Teradata, MongoDB, Cassandra DB, MySQL, Oracle, and SQL Server.
- Experience in Creating Teradata SQL scripts using OLAP functions like rank and rank over to improve the query performance while pulling the data from large tables.
- Strong Experience in working with Databases like Teradata and proficiency in writing complex SQL, PL/SQL for creating tables, views, indexes, stored procedures, and functions.
- Knowledge and experience with Continuous Integration and Continuous Deployment using containerization technologies like Docker and Jenkins.
- Excellent working experience in Agile/Scrum development and Waterfall project execution methodologies.
TECHNICAL SKILLS
Big Data Technologies: Hadoop, MapReduce, Spark, HDFS, Sqoop, YARN, Oozie, Hive, Impala, Zookeeper, Apache Flume, Apache Airflow, Cloudera, HBase
Programming Languages: Python, PL/SQL, SQL, Scala, PowerShell, C, C++, T-SQL
Cloud Services: Azure Data Lake Storage Gen 2, Azure Data Factory, Blob storage, Azure SQL DB, Databricks, Azure Event Hubs, AWS RDS, Amazon SQS, Amazon S3, AWS EMR, Lambda, AWS SNS, Big Query, Data Proc, Data Flow.
Databases: MySQL, SQL Server, Oracle, MS Access, Teradata, and Snowflake
NoSQL Data Bases: MongoDB, Cassandra DB, HBase
Monitoring tool: Apache Airflow
Visualization & ETL tools: Tableau, Informatica, Talend, SSIS, and SSRS
Version Control & Containerization tools: Jenkins, Git, and SVN
Operating Systems: Unix, Linux, Windows, Mac OS
PROFESSIONAL EXPERIENCE
Sr. Data Engineer
Confidential, Florida
Responsibilities:
- Worked with business/user groups for gathering the requirements and working on the creation and development of pipelines.
- Migrated applications from Cassandra DB to Azure Data Lake Storage Gen 2 using Azure Data Factory, created tables, and loading and analyzed data in the Azure cloud.
- Worked on creating Azure Data Factory and managing policies for Data Factory and Utilized Blob storage for storage and backup on Azure.
- Worked on developing the process and ingested the data in Azure cloud from web service and load it to Azure SQL DB.
- Worked with Spark applications in Python for developing the distributed environment to load high volume files using PySpark with different schema into PySpark Data frames and process them to reload into Azure SQL DB tables.
- Used Sales force for the extract, load, transform process that helps in extraction for complex models.
- Installed Kafka on Hadoop cluster and configured producer and consumer in java to establish connection from source to HDFS with popular hash tags.
- Used Ruby for Data Processing and web scrapping.
- Designed and developed the pipelines using Databricks and automated the pipelines for the ETL processes and further maintenance of the workloads in the process.
- Worked on creating ETL packages using SSIS to extract data from various data sources like Access database, Excel spreadsheet, and flat files, and maintain the data using SQL Server.
- Worked with ETL operations in Azure Databricks by connecting to different relational databases using Kafka and used Informatica for creating, executing, and monitoring sessions and workflows.
- Worked on automating data ingestion into the Lakehouse and transformed the data, used Apache Spark for leveraging the data, and stored the data in Delta Lake.
- Ensured data quality and integrity of the data using Azure SQL Database and automated ETL deployment and operationalization.
- Used Databricks, Scala, and Spark for creating the data workflows and capturing the data from Delta tables in Delta Lakes.
- Performed Streaming of pipelines using Azure Event Hubs and Stream Analytics to analyze the data from the data-driven workflows.
- Used Sales Force for enhanced security while transferring complex data.
- Worked with Delta Lakes for consistent unification of Streaming, processed the data, and worked on ACID transactions using Apache Spark.
- Worked with Azure Blob Storage and developed the framework for the implementation of the huge volume of data and the system files.
- Implemented of distributed stream processing platform with low latency and seamless integration, with data and analytics services inside and outside Azure to build your complete big data pipeline.
- Worked with PowerShell scripting for maintaining and configuring the data. Automated and validated the data using Apache Airflow.
- Worked on optimization of Hive queries using best practices and right parameters and using Hadoop, YARN, Python, and PySpark.
- Used Sqoop to extract the data from Teradata into HDFS and export the patterns analyzed back to Teradata.
- Worked on Kafka to bring the data from data sources and keep it in HDFS systems for filtering.
- Used Accumulators and Broadcast variables to tune the Spark applications and to monitor the created analytics and jobs.
- Used Sales force for the CRM data in a single data store like warehouse for consistency and standardidation.
- Tracked Hadoop cluster job performance and capacity planning and tuning Hadoop performance for high availability and Hadoop cluster recovery.
- Worked with Tableau for generating reports and created Tableau dashboards, pie charts, and heat maps according to the business requirements.
- Worked with all phases of Software Development Life Cycle and used Agile methodology for development.
Environment: Python, SQL, Sales force,Cassandra DB, Azure Data Lake Storage Gen 2, Azure Data Factory, Azure SQL DB, Spark, Databricks, SSIS, SQL Server, Kafka, Informatica, Apache Spark, Delta Lake, Azure Event Hubs, Stream Analytics, Azure Blob Storage, PowerShell, Apache Airflow, Hadoop, YARN, PySpark, Hive, Teradata, Sqoop, HDFS, Spark, Agile.
Data Engineer
Confidential, Charlotte
Responsibilities:
- Worked in complete Software Development Life Cycle (SDLC) process by analyzing business requirements and understanding the functional workflow of information from source systems to destination systems.
- Utilizing analytical, statistical, and programming skills to collect, analyze and interpret large data sets to develop data-driven and technical solutions to difficult business problems using tools such as SQL, and Python.
- Worked on designing AWS EC2 instance architecture to meet high availability application architecture and security parameters.
- Created AWS S3 buckets and also managed policies for S3 buckets and Utilized S3 buckets and Glacier for storage and backup.
- Worked on Hadoop cluster and data querying tools to store and retrieve data from the stored databases.
- Worked with different file formats like Parquet files and also Impala using PySpark for accessing the data, and performed Spark Streaming with RDDs and Data Frames.
- Performed the aggregation of log data from different servers and used them in downstream systems for analytics using Apache Kafka.
- Worked on designing and developing the SSIS Packages to import and export data from MS Excel, SQL Server, and Flat files.
- Worked on Data Integration for extracting, transforming, and loading processes for the designed packages.
- Designed and deployed automated ETL workflows using AWS lambda, organized and cleansed the data in S3 buckets using AWS Glue, and processed the data using Amazon Redshift.
- Worked within the ETL architecture enhancements to increase the performance using query optimizer.
- Implemented the data that is extracted using Spark, Hive, and large data sets using HDFS.
- Worked on Streaming data transfer, data from different data sources into HDFS, No SQL databases.
- Created ETL Mapping with Talend Integration Suite to pull data from Source, apply transformations, and load data into the target database.
- Worked on scripting with Python in Spark for transforming the data from various files like Text files, CSV and JSON.
- Loaded the data from different relational databases like MySQL and Teradata using Sqoop jobs.
- Worked on processing the data and testing using Spark SQL and on real-time processing by Spark Streaming and Kafka using Python.
- Scripted using Python and PowerShell for setting up baselines, branching, merging, and automation processes across the process using GIT.
- Worked with the implementation of the ETL architecture for enhancing the data and optimized workflows by building DAGs in Apache Airflow to schedule the ETL jobs and additional components in Apache Airflow like Pool, Executors, and multi-node functionality.
- Used various Transformations in SSIS Dataflow, Control Flow using for loop Containers and Fuzzy.
- Worked on creating SSIS packages for Data Conversion using data conversion transformation and producing the advanced extensible reports using SQL Server Reporting Services.
Environment: Python, SQL, AWS EC2, AWS S3 buckets, Hadoop, PySpark, AWS lambda, AWS Glue, Amazon Redshift, Spark Streaming, Apache Kafka, SSIS, Informatica, ETL, Hive, HDFS, NoSQL, Talend, MySQL, Teradata, Sqoop, PowerShell, GIT, Apache Airflow.
Data Engineer
Confidential
Responsibilities:
- Experience in building and architecting multiple Data pipelines, end to end EL and ELT process for Data ingestion and transformation in GCP and coordinate task among the team.
- Implemented and Managed ET solutions and automating operational processes
- Design and develop ET integration patterns using Python on Spark
- Develop framework for converting existing PowerCenter mappings and to PySpark (Python and Spark) Jobs.
- Build data pipelines in airflow in GCP for ET related jobs using different airflow operators.
- Used Stitch ETL tools to integrate data into the central data warehouse.
- Experience in GCP Dataproc, GCS, Cloud functions, Data prep, Data Studio and Big Query.
- Implemented Spark RDD transformations to map business analysis and apply actions on top of Transformations.
- Design star schema in Big Query.
- Worked on creating various types of indexes on different collections to get good performance in Mongo database.
- Monitoring Big query, Dataproc and cloud Data flow jobs via Stack driver for all the environments.
- Used Agile for the continuous model deployment.
- Worked with Google data catalog and other google cloud APIs for monitoring, query and billing related analysis for big query usage.
- Knowledge about cloud data flow and Apache beam.
- Write Scala program for spark transformation in Dataproc.
- Used Snowflake for the Data Storage, processing which is easier and faster to use.
- Write a Python program to maintain raw file archival in GCS bucket.
- Write Scala program for spark transformation in Dataproc.
- Worked with google data catalog and other google cloud APl's for monitoring, query, and billing related analysis for Big Query usage.
- Used Airflow to manage task scheduling, progress, and success status using DAG graphs.
- Created Big Query authorized views for row level security or exposing the data to other teams.
- Integrated services like GitHub, Jenkins to create a deployment pipeline.
- Implemented new projects builds framework using Jenkins as build framework tools.
Environment: T-SQL, PL/SQL, Google Cloud, Python, Big query, Dataflow, Dataproc, Dataprep, Data Studio, Bigtable Stitch ETL, PySpark, Snowflake, MySQL, Airflow, Shell Scripts, Mongo DB, GIT, Apache, Spark, Docker
