Senior Data Engineer Resume
Glen Allen, VA
SUMMARY
- 7+ years experienced certified Data Engineer adept at creating data set processes, maintaining scalable data pipelines, building API Integration, and accelerating ETL performances. Experience in the waterfall / Agile and Kanban methodologies, specialized in Hadoop/Big Data ecosystem, Data Acquisition, Ingestion, Modeling, Storage Analysis, Integration, and Data Processing, Cloud Engineering, Data Warehousing.
- Excellent knowledge on Hadoop architecture and various components such as HDFS Job Tracker, Task Tracker, Name Node (Master Node), Data Node (Slave Node) and MapReduce programming paradigm.
- Hands on experience in using Hadoop ecosystem components like Hive, Pig, Sqoop, HBase, Cassandra, Spark, Spark Streaming, Spark SQL, Oozie, Zookeeper, Flume, Kafka, Nifi, MapReduce, Yarn, and Scala.
- Hands on experience about architecture and components of Spark, and efficient in working with Hadoop Core, Spark SQL, Spark streaming and expertise in building PySpark and Spark - Scala applications for Interactive analysis, batch processing and stream processing.
- Acquired profound knowledge in developing production ready Spark applications utilizing Spark Core, Spark Streaming, Spark SQL, Data Frames, Datasets and Spark-ML.
- Designed scalable Scala Web Architecture hosting reports for the entire application. Wrote entities in Scala and Java along with named queries to interact with database.
- Mastered development of applications/tools using Python and worked on several python packages like NumPy, Pandas, Matplotlib, Seaborn, Scikit-learn, SciPy, Pytables etc.
- Experience with Design, Code, Debug Operations, Reporting, Data analysis and web applications utilizing Python and Python Web Framework Django.
- Used traditional databases like MS SQL Server, MySQL, and Oracle to move data into Hadoop by using HDFS and performed Data Profiling and Data Analysis using SQL on different extracts.
- Support and management of NoSQL database Install, configure, administer, and support multiple NoSQL instances perform database maintenance and trouble shooting.
- Hands on experience on Unified Data Analytics with Databricks, Databricks Workspace User Interface, Managing Databricks Notebooks, Delta Lake with Python, Delta Lake with Spark SQL.
- Hands on experience in Data Migration, Data Profiling, Data Cleansing, Transformation, Integration, Data Import, and Data Export using multiple ETL tools such as Informatica, SSIS.
- Expertise in Informatica PowerCenterin all stages of design, development, and implementation of datamappings,mapplets,sessionsusing Informatica PowerCenterIDQ,Oracle,SQL,andUnix
- Expertise in developing multiple Kafka Producers and Consumers as per the software requirement specifications. Have experience in developing a knowledge pipeline using Kafka to store data into HDFS.
- Created multi-tier java based multiple web services to read data from MongoDB and developed core java classes for exceptions, utility classes, business delegate and test cases.
- Expertise in loading data into Snowflake DB in the cloud from various sources. Validated the data feed from the source system to Snowflake DW Cloud platform.
- Drew upon full range of Tableau platform technologies to design and implement proof of concept solutions and create advanced BI visualizations. Created Tableau scorecards, dashboards using stack bars, bar graphs, scattered plots, geographical maps, Gantt charts using show me functionality.
- Experienced in working with Amazon Web Services (AWS) using EC2 for computing and S3 for storage.
- Developed data transition programs from DynamoDB to AWS Redshift (ETL Process) using AWS Lambda by creating functions in Python for the certain events based on use cases.
- Used Azure Synapse to manage processing workloads and served data for BI and prediction needs.Knowledge in setting Up AWS and Microsoft Azure with Databricks, Databricks Workspace for Business Analytics, Manage Clusters in Databricks, Managing the Machine Learning Lifecycle
- Managed resources and scheduling across the cluster usingAzure Kubernetes Service, and worked onAzure Data Factory, Azure Databricks, Azure Data Lake, SQL API, and Mongo API and integrated data from MongoDB, MS SQL, and cloud (Blob, Azure SQL DB).
- Knowledge on Azure data factory with different triggers and Azure storage accounts and migrating existing infrastructure and configures and Manage Azure site Rec.
- Worked on Google Cloud Platform (GCP) services like cloud storage, cloud SQL, stack driver monitoring.
- Developed ETL solution for GCP Migration using GCP Dataflow, GCP Composer, Apache Airflow and GCP BigQuery. Hands on experience in GCS bucket, G- cloud function, cloud dataflow, Pub/sub cloud shell, GSUTIL, BQ command line utilities, Data Proc, Stack driver.
- Managed SCM processes with tools like GitHub, Bitbucket, GitLab, and defined branching strategies, pull requests, solved merge conflicts, and worked setting up centralized version control system such as Subversion (SVN) and distributed version control system such as GIT.
- Expertise in installation, integration, and configuration of CI/CD tools like GitLab CI, GitHub actions, Azure DevOps and Jenkins including installation of Jenkins plugins.
- Involved in loading of data into Terradata from legacy systems and flat files using complex scripts. Teradata performance tuning via Explain, PPI, AJI, Indices, collect statistics or rewriting of the code.
- Worked with various monitoring tools like AppDynamics, Datadog, ELK, Dynatrace etc. Deployed monitoring modules to new system components sing AppDynamics.
- Expert in installing and using Splunk apps for UNIX and LINUX. Implemented workflow actions to drive troubleshooting across multiple event types in Splunk.
- Customized and branded Jira (server and cloud) and used Jira bash shell scripting. Documented the connection between Jira Align and Jira using align connectors.
TECHNICAL SKILLS
Frame Works Tool Kits: Hadoop/Big Data
Frame Works: HDFS, Apache NIFI, Map Reduce, Sqoop, Flume, Pig, Hive, Oozie, Impala, Zookeeper, Ambari, Storm.
Cloud Technologies: AWS, GCP, Azure
Programing Languages: Python, Scala, Unix Shell scripting, Spark SQL, HiveQL, R, C, C++
Databases: MySQL, MS-SQL Server, NoSQL (HBase, MongoDB), PostgreSQL
Reporting/ETL Tools: Airflow, Kafka, Informatica, Pentaho, Talend, Tableau, SSIS
Operating Systems: Windows, MacOS, Linux (Ubuntu), CentOS
Bigdata distribution: Cloudera, Hortonworks, Amazon EMR
IDE’S: Hue, Eclipse, IntelliJ IDEA, NetBeans, MS Office Suite.
Version Control Tools: Git, GitHub, Bitbucket, GitLab
PROFESSIONAL EXPERIENCE
Confidential, Glen Allen, VA
Senior Data Engineer
Responsibilities:
- Experienced in ET concepts, building ETL solutions and Data modeling and worked on architecting the ETL transformation layers and writing spark jobs to do the processing.
- Experienced on loading and transforming of large sets of structured, semi structured, and unstructured data and Developed a Spark Streaming module for consumption of Avro messages from Kafka.
- Aggregated daily sales team updates to send report to executives and to organize jobs running on Spark clusters and Loaded application analytics data into data Warehouse in regular intervals of time.
- Developed pyspark programs and created the data frames and worked on transformations. Worked on data processing and transformations and actions in spark by using Python (Pyspark) language.
- Optimized the Pyspark jobs to run on Kubernetes Cluster for faster data processing. Worked on reading and writing multiple data formats like JSON, ORC, Parquet on HDFS using Pyspark.
- Implemented in setting up standards and processes for Hadoop based application design and implementation. Good Knowledge on Hadoop Cluster architecture and monitoring the cluster.
- Expertise in analysis, design and development using Hadoop, HDFS, Hortonworks, MapReduce, and Hadoop Ecosystem (Pig, Hive, Impala and Spark, Scala), Java and J2EE.
- Optimize Map Reduce Programs using combiners, partitioners, and custom counters for delivering the best results and Written Shell scripts to monitor the health check of Hadoop daemon services and respond accordingly to any warning or failure conditions and involved in Hadoop cluster task like Adding and Removing Nodes without any effect to running jobs and data.
- Designed and implemented configurable data delivery pipeline for scheduled updates to customer facing data stores built with Python and Proficient in Machine Learning techniques (Tree based algorithms like Decision Trees, Random Forest, Linear/Logistic Regressors) and Statistical Modeling.
- Involved in building database Model, APIs and Views utilizingPython, to build an interactive web-based solution and developed Merge jobs inPythonto extract and load data into MySQL database.
- Developed SQL Scripts for automation purpose and encapsulated frequently executed SQL Statements in stored procedure to reduce query execution time.
- Automated DBA activities such as backups and upgraded existing PostgreSQL databases and maintained and optimized PostgreSQL database application performance through tuning.
- Generated PostgreSQL database reports such as financial data statements and user data and administered and troubleshoot PostgreSQL databases for critical problems.
- Designed several dashboards using Tableau Desktop, Tableau Server, Tableau Reader, Tableau Online.
- Conducted data cleansing for unstructured dataset by applyingInformatica Data Qualityto identify potential errors and improvedata integrityanddata quality.
- Used most of the Transformations available in Informatica -Source Qualifier, Filter, Router, Lookup (Connected & Un Connected), Expression, Update Strategy, Transaction Control and Sequence Generator.
- Involved inInformatica upgrade process and testing the whole existing Informatica flowin new upgrade environment.
- Implemented Apache Airflow for authoring, scheduling, and monitoring Data Pipelines and Designed several DAGs for automating ETL pipelines and Performed data extraction, transformation, loading, and integration in data warehouse, operational data stores and master data management.
- Hands on experience withClouderaand multi cluster nodes onHortonworks Sandbox. CreatedODBCconnection throughSqoopbetweenHortonworksandSQL Server.
- Leveraged cloud and GPU computing technologies for automated machine learning and analytics pipelines on cloud platforms such as GCP.
- Designed & build infrastructure for the Google Cloud environment from scratch along with fact dimensional modeling (Star schema, Snowflake schema), transactional modeling and SC.
- Used SSH tunnel to Google DataProc to access to yarn manager to monitor spark jobs. Monitoring Bigquery, Dataproc and cloud Data flow jobs via Stackdriver for all the environments.
- Monitoring Bigquery, Dataproc and cloud Data flow jobs via Stackdriver for all the environments.
- Used Google cloud services like GKE and BigQuery to easily integrate with existing apps. Implemented scripts that Load google BigQuery data and run queries to export data.
- Used Zookeeper to provide coordination services to the cluster and Created Hive queries that helped analysts spot emerging trends by comparing fresh data with reference tables and historical metrics.
- Involved in migrating tables from RDBMS into Hive tables using SQOOP and later generated data visualizations using Tableau.
- Responsible for deploying Ab Initio graphs and running them through the Co-operating systems MP shell command language and responsible for automating the ETL process through scheduling
- Worked on improving the performance of Ab Initio graphs by using Various Ab Initio performance techniques like using lookups In-Memory Joins and rollups to speed up various Ab Initio Graphs
- Process large volume of data and skills in parallel execution of process using Abinitio functionality and Designed implemented Spark jobs to support distributed data processing.
- Used Azure Data Factory extensively for ingesting data from disparate source systems. Used Azure Data Factory as an orchestration tool for integrating data from upstream to downstream systems.
- Analyzed the data flow from different sources to target to provide the corresponding design Architecture in Azure environment.
- Created Application Interface Document for the downstream to create new interface to transfer and receive the files through Azure Data Share
Environment: Hadoop, Kafka, Azure, Gcp, Dataproc, MYSQL, Postgres, SQL Server, Python, Scala, Spark, Hive, Spark -SQL, Tableau.
Confidential, Irvine, CA
Data Engineer
Responsibilities:
- Managed and scheduled Jobs on a Hadoop Cluster using UC4(Confidential preoperatory scheduling tool) workflows, monitored and managed Hadoop cluster through Hortonworks (HDP) distribution.
- Analyzed substantial data sets by running Hive queries and Pig scripts and Managed and reviewed Hadoop and Base log files along with the Experience in creating tables, dropping, and altered at run time without blocking updates and queries using Spark and Hive.
- Performed data transformations by writing MapReduce and Pig scripts as per business requirements and Implemented Map Reduce programs to handle semi structured data, unstructured data like xml, json, Avro data files and sequence files for log files.
- Designed and deployed Hadoop cluster and different big data analytic tools including Pig, Hive, Flume, HBase and Sqoop. Loaded the DRs from relational DB using Sqoop to Hadoop cluster.
- Developed Spark SQL logics which mimics Teradata ET logics and point the output Delta back to Newly Created Hive Tables and as well the existing TERADATA Dimensions, Facts, and Aggregated Tables.
- Imported data from Abinitio LDR and into Spark RDD and performed transformations, actions on RDD's. Developed spark job with partitioned RDDB (like hash, range, custom) for faster processing.
- Implementing quality checks and transformations using Spark and Developed simple and complex MapReduce programs in Hive, Pig and Python for Data Analysis on different data formats.
- Implemented Spark using Scala and utilizingData framesand Spark SQL API for faster processing of data. Expertized in implementing Spark usingScalaandSpark SQLfor faster testing and processing of data responsible to manage data from different sources.
- Developed multiple POCs using Scala and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata.
- Developed various Python scripts to find vulnerabilities with SQL Queries by doing SQL injection, permission checks and analysis and experienced in Kerberos authentication to establish a more secure network communication on the cluster.
- Wrote Spark Applications in Scala and Python along with Used Spark SQL to handle structured data in Hive and Imported semi-structured data from Avro files using Pig to make serialization faster
- Converted Hive/SQL queries into Spark transformations using Spark DD, Scala and Python and Experienced in connecting Avro Sink ports directly to Spark Streaming for analyzation of weblogs.
- Involved in making Hive tables, stacking information, composing hive inquiries, producing segments and basins for enhancement.
- Configured various views in Yarn Queue manager. Involved in review of functional and non-functional requirements and Indexed documents using Elastic search and Responsible for using Flume sink to remove the date from Flume channel and deposit in No-SQL database like MongoDB
- Involved in loading data from UNIX file system and FTP to HDFS and Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig. Loaded JSON-Styled documents in NoSQL database like MongoDB and deployed the data in cloud service Amazon Redshift.
- Responsible for developing data pipeline with AWS to extract the data from weblogs and store in Amazon EMR. Deployed applications on AWS by using Elastic Beanstalk.
- Orchestrated the digital transformation at the application and infrastructure levels to deliver meaningful transformations and unleash the full capabilities of modern cloud service providers (AWS, Azure, VMware) and cloud platforms (PCF) to increase time-to-value of products for organizations.
- Managed security groups on AWS, focusing on high availability, fault-tolerance, and auto-scaling using terraform templates, along with continuous integration and continuous deployment with AWS Lambda and AWS code pipeline.
- Strong understanding of AWS components such as EC2 and S3 and Performed Data Migration to GCP and Responsible for data services and data movement infrastructures.
- Provisioned the highly available EC2 Instances using Terraform and cloud formation and wrote new plugins to support new functionality in Terraform.
- Construct the AWS data pipelines using VPC, EC2, S3, Auto Scaling Groups (ASG), EBS, Snowflake, IAM, CloudFormation, Route 53, CloudWatch, CloudFront, CloudTrail.
- Used the AWS -CLI to suspend on Aws Lambda function used AWS CLI to automate backup of ephemeral data stores to S3 buckets EBS. Created functions and assigned roles in AWS Lambda to run python scripts, and AWS Lambda using java to perform event driven processing.
- Developed data transition programs from DynamoDB to AWS Redshift (ETL Process) using AWS Lambda by creating functions in Python for the certain events based on use cases.
- Used Power BI to create self-service BI capabilities and use tabular models. Performed time series analysis using DAX expression in Power BI and Worked with both live and import data into Power BI for creating reports.
- Designed and developed a new solution to process the RT data by using Azure stream analytics, Azure Event Hub, and Service Bus Queue. Created Linked service to land the data from SFTP location to Azure Data Lake.
- Knowledge of designing SCD Type 1andType 2in Informatica PowerCenter using multiple transformations. ETL created by multiple Informatica transformations (Source Qualifier, Lookup, Router, Update Strategy) were utilized to create SCD type mappings
- Performed end-to-end delivery of pyspark ETL pipelines on Azure-databricks to perform the transformation of data orchestrated via Azure Data Factory (ADF) scheduled through Azure automation accounts and trigger them using Tidal Schedular.
- Experience in developing Spark applications using Spark-SQL inDatabricksfor data extraction, transformation, and aggregation from multiple file formats.
- Understanding of API management with security, integration in Azure Cloud and Configuring backup and recovery of Azure Resources.
- Scheduled reports, report services documents and dashboards using Schedule Manager. Knowledge in developing reports in Visual Studio and deploying them to Report Manager.
- Highly experienced in ETL tool Ab Initio in using GDE Designer. Very good working experience with all the Ab Initio components.
- Expertise and well versed with various Ab Initio Transform, Partition, Departition, Dataset Database components, sort, validate and compress.
Environment: Hadoop, Python, SQL, AWS, Azure, Power BI, Visual Studio, Informatica, Abinitio, etc.
Confidential
Data Engineer
Responsibilities:
- Developed Spark scripts using Python on Azure HDInsight for Data Aggregation, Validation and verified its performance over MR jobs.
- Utilized Azure HDInsight to monitor and manage the Hadoop Cluster.
- Automated Sqoop incremental imports by using Sqoop jobs and automated jobs using Oozie.
- Worked on various compression and file formats like Avro, Parquet, and Text formats.
- Responsible for writing Hive Queries for analyzing data in Hive warehouse using HQL.
- Experience with CDH distribution and used Cloudera Manager for installation and management of single-node and multi-node Hadoop cluster (CDH3, CDH4 & CDH5).
- Created significant improvement in the execution times to Cluster Migration from Hortonworks distribution to Cloudera. Managing and monitoring the Hadoop cluster through Cloudera Manager
- Consumed Extensible Markup Language (XML) messages using Kafka and processed the xml file using Spark Streaming to capture User Interface (UI) updates.
- Designed multiple Python packages that were used within a large ETL process used to load 2TB of data from an existing Oracle database into a new PostgreSQL cluster.
- Configured PostgreSQL Streaming Replication and Pgpool for load balancing. Managing the monitoring tools for better performance like PgBadger, Kibana, Graphana, and Nagios.
- Responsible for all backup, recovery, and upgrading of all the PostgreSQL databases. Monitoring databases to optimize database performance and diagnosing any issues.
- Proactive in updating the security patches to database, provided by PostgreSQL open-source community, Setup, and maintenance of postgres master-slave clusters utilizing streaming replication.
- Developed MapReduce/Spark Python modules for machine learning predictive analytics in Hadoop.
- Leveraged ETL methods for ETL solutions and data warehouse tools for reporting and analysis. Used CSV Excel Storage to parse with different delimiters in PIG.
- Performed advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark using Scala.
- Experienced in Data Stewardship, Master Data Management (MDM), MDM - Member domain (Patient), MDM Provider domain (Prescriber), Member Enrollment, Claims, Data Governance aspects of Data Management in the Healthcare domain.
- Expertise in Master Data Management concepts, Methodologies, and ability to apply this knowledge in building master data management solutions using Informatica MDM and IDD tools.
Environment: s: Python, Azure, Hadoop, Sqoop, Oozie, Avro, Parquet, Hive Queries, Kafka, SQL, Spark, Scala.
Confidential
Data Engineer Associate
Responsibilities:
- ImplementedSparkusing Scala and Spark SQL for faster testing and processing of data.
- Extensively usedPre-SQL and Post-SQL scriptsfor loading the data into the targets according to the requirement.
- Worked closely with multi-functional teams and collaborated on efficient solutions while setting future targets and formulating enhanced strategies & plans
- Competently utilized numerous technical skills including R studio, Python and My SQL Workbench for comprehensive data analysis and visualization
- Created a TensorFlow session which is used to run the neural network as well as validate the accuracy of the model on the validation set.
- Executed multiple Spark SQL queries after forming the Database to gather required data.
Environment: Python, R studio, My SQL Workbench, Scala, TensorFlow, Spark SQL Queries, etc.
