We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

SUMMARY

  • 7 years of overall experience in working as a Data Engineer, ETL Developer and Hadoop Developer comprising planning, improvement, and execution of Data models at an enterprise level application.
  • Experience in application development, implementation, deployment, and maintenance using Hadoop and Spark - based technologies like Cloudera, Hortonworks, Amazon EMR, Azure HDInsight.
  • A Data Science enthusiast with strong Problem solving, Debugging, and Analytical capabilities, who actively engage in understanding and delivering to business requirements.
  • Experience in Systems Analysis, Design, Development of various Client/Server and Internet Applications using Java, JSP, jQuery, Struts, Spring, Hibernate, JDBC Servlets, Java Beans, HTML, XML, JavaScript.
  • Ample work experience in Big-Data ecosystem - Hadoop (HDFS, MapReduce, Yarn), Spark, Kafka, Hive, Impala, HBase, Sqoop, Pig, Airflow, Oozie, Zookeeper, Ambari, Flume.
  • Good knowledge of Hadoop cluster architecture and its key concepts - Distributed file systems, Parallel processing, High availability, fault tolerance, and Scalability.
  • Complete knowledge of Hadoop architecture and Daemons of Hadoop clusters, which include Name node, Data node, Resource manager, Node Manager, and Job history server.
  • Expertise in developing Spark applications for interactive analysis, batch processing and stream processing, using programming languages like PySpark, Scala and Java.
  • Advanced knowledge in Hadoop based Data Warehouse (HIVE) and database connectivity (SQOOP).
  • Ample experience using Sqoop to ingest data from RDBMS - Oracle, MS SQL Server, Teradata, PostgreSQL, and MySQL.
  • Experience in working with various streaming ingest services with Batch and Real-time processing using Spark Streaming, Kafka, Confluent, Storm, Flume, and Sqoop.
  • Proficient in using Spark API for streaming real-time data, staging, cleaning, applying transformations, and preparing data for machine learning needs.
  • Experience in developing end-to-end ETL pipelines using Snowflake, Alteryx, and Apache NiFi for both relational and non-relational databases (SQL and NoSQL).
  • Strong working experience on NoSQL databases and their integration with the Hadoop cluster - HBase, Cassandra, MongoDB, DynamoDB, and Cosmos DB.
  • Experience with AWS cloud services to develop cloud-based pipelines and Spark applications using EMR, LAMBDA and Redshift.
  • Extensive knowledge in working with Amazon EC2 to provide a solution for computing, query processing, and storage across a wide range of applications.
  • Expertise in using AWS S3 to stage data and to support data transfer and data archival. Experience in using AWS Redshift for large scale data migrations using AWS DMS and implementing CDC (change data capture).
  • Strong experience in developing LAMBDA functions using Python to automate data ingestion and tasks.
  • Working knowledge of Azure cloud components (HDInsight, Databricks, Data Lake, Blob Storage, Data Factory, Storage Explorer, SQL DB, SQL DWH, Cosmos DB).
  • Experienced in building data pipelines using Azure Data Factory, Azure Databricks, and loading data to Azure Data Lake, Azure SQL Database, Azure SQL Data Warehouse, and controlling database access.
  • Extensive experience with Azure services like HDInsight, Stream Analytics, Active Directory, Blob Storage, Cosmos DB, and Storage Explorer.
  • Good knowledge in understanding the security requirements and implementation using Azure Active Directory, Sentry, Ranger, and Kerberos for authentication and authorizing resources.
  • Experience in all phases of Data Warehouse development like requirements gathering, design, development, implementation, testing, and documentation.
  • Extensive knowledge of Dimensional Data Modelling with Star Schema and Snowflake for FACT and Dimensions Tables using Analysis Services.
  • Good experience in the development of Bash scripting, T-SQL, and PL/SQL Scripts.
  • Sound knowledge in developing highly scalable and resilient Restful APIs, ETL solutions, and third-party platform integrations as part of Enterprise Site platform.
  • Experience in designing interactive dashboards, reports, performing ad-hoc analysis and visualizations using Tableau, Power BI, Arcadia, and Matplotlib.
  • Experience in implementing pipelines using ELK (Elasticsearch, Logstash, Kibana) and developing stream processes using Apache Kafka.
  • Sound knowledge and experience in programming languages like JAVA, Python, Scala.
  • Experience in using various IDEs like Eclipse, IntelliJ, and repositories SVN and Git version control systems.
  • A team player with strong communication, Interpersonal, problem-solving, and debugging skills. Ability to quickly adapt to new environments and technologies.
  • Successfully working in a fast-paced environment, both independently and in a collaborative way. Expertise in complex troubleshooting, root-cause analysis, and solution development.

TECHNICAL SKILLS

Big Data Ecosystem: HDFS, Yarn, MapReduce, Spark, Kafka, Kafka Connect, Hive, Airflow, Stream Sets, Sqoop, HBase, Flume, Pig, Ambari, Oozie, Zookeeper, Nifi, Sentry

Hadoop Distributions: Apache Hadoop 2.x/1.x, Cloudera CDP, Hortonworks HDP

Cloud Environment: Amazon Web Services (AWS), Microsoft Azure

Databases: MySQL, Oracle, Teradata, MS SQL SERVER, PostgreSQL, DB2

NoSQL Database: Cassandra, MongoDB, DynamoDB, Cosmos DB

AWS: EC2, EMR, S3, Redshift, EMR, Lambda, Kinesis Glue, Data Pipeline

Microsoft Azure: Databricks, Data Lake, Blob Storage, Azure Data Factory, SQL Database, SQL Data Warehouse, Cosmos DB, Azure Active Directory

Operating systems: Linux, Unix, Windows 10, Windows 8, Windows 7, Windows Server 2008/2003, Mac OS

Reporting Tools/ETL Tools: Informatica, Talend, SSIS, SSRS, SSAS, ER Studio, Tableau, Power BI, Arcadia, Data stage, Pentaho

Programming Languages: Python (Pandas, SciPy, NumPy, Scikit-Learn, Stats Models, Matplotlib, Plotly, Seaborn, Keras, TensorFlow, Pytorch), PySpark, T-SQL/SQL, PL/SQL, HiveQL, Scala, UNIX Shell Scripting.

Version Control: Git, SVN, Bitbucket

PROFESSIONAL EXPERIENCE

Data Engineer

Confidential

Responsibilities:

  • Worked on Apache Spark data processing project to process data from RDBMS and several data streaming sources and developed Spark applications using Python on AWS EMR.
  • Designed and deployed multi-tier applications leveraging AWS services like (EC2, Route 53, S3, RDS, DynamoDB) focusing on high availability, fault tolerance, and auto-scaling in AWS Cloud Formation.
  • Configured and launched AWS EC2 instances to execute Spark jobs on AWS Elastic Map Reduce (EMR).
  • Performed data transformations using Spark Data Frames, Spark SQL, Spark File formats, Spark RDDs.
  • Transformed data from different files (Text, CSV, JSON) using Python scripts in Spark.
  • Loaded data from various sources like RDBMS (MySQL, Teradata) using Sqoop jobs.
  • Handled JSON datasets by writing custom Python functions to parse through JSON data using Spark.
  • Developed a pre-processing job using Spark Data Frames to flatten JSON documents to flat files.
  • Improved performance of cluster by optimizing existing algorithms using Spark.
  • Performed wide, narrow transformations, actions like filter, Lookup, Join, count, etc. on Spark Data Frames.
  • Worked with Parquet files and Impala using PySpark, and Spark Streaming with RDDs and Data Frames.
  • Aggregated logs data from various servers and made them available in downstream systems for analytics by using Apache Kafka. Improved Kafka performance and implemented security.
  • Developed batch and streaming processing apps using Spark APIs for functional pipeline requirements.
  • Automated data storage from streaming sources to AWS data lakes like S3, Redshift and RDS by configuring AWS Kinesis (Data Firehose).
  • Performed analytics using real time integration capabilities of AWS Kinesis (Data Streams) on streamed data
  • Cleaned and handled missing values in data using Python by backward-forward filling methods and applied Feature engineering, normalize and label encoding techniques using Python Scikit-learn pre-processing.
  • Developed data pipelines from on premises to cloud using AWS services like DMS, S3, EC2, EMR and Redshift.
  • Configured and built AWS Lambda functions to automate data ingestion and terminate resources not in use.
  • Writing Python Applications which runs EMR cluster that fetches data from the S3/one lake location and queue it in the Amazon SQS (simple Queue Services) queue.
  • Stored data into various tiers of AWS S3 based on business requirements and frequency of data access.
  • Performed reporting analytics on data from AWS stack by connecting it to BI tools (Tableau, Power Bi).
  • Imported data from AWS S3 intoSpark RDD performed transformations and actions on RDD's.
  • Worked with database administrating team on SQL optimization for databases like Oracle, MySQL, MS SQL.
  • Assisted in configuring and implemented MongoDB cluster nodes on AWS EC2 instances.
  • Identified executor failures, data skewness, and runtime issues by monitoring Spark apps through Spark UI.
  • Ensured database performance in production by stress testing AWS EC2 of MongoDB environments.
  • Automated deployments and routine tasks using UNIX Shell Scripting.
  • Worked in an agile environment to implement projects and enhancements with weekly SCRUMs.

Environment: DynamoDB, AWS Glue, AWS Kinesis, AWS CloudWatch, MapReduce, PySpark, Python, AWS(Lambda, Step Functions, SQS, Event Bridge, Athena), AWS Redshift, JSON, XML, AVRO, EMR, AWS S3,Unix/Linux Shell Scripting, Hadoop, SQL Server, Oracle.

Azure Data Engineer

Confidential

RESPONSIBILITIES:

  • Integrated Big Data/Hadoop Distribution Frameworks: Zookeeper, Yarn, Spark, Scala, NiFi, and others with Advanced with Azure cloud platforms (hdinsight, datalake, databricks, Blob Storage, Data Factory, Synapse, SQL, SQL DB, DWH and Data Storage Explorer).
  • Acted aa an admin for the Azure SQL Database, Azure Analysis Service, Azure SQL Data warehouse, Azure Data Factory, and Azure SQL Data warehouse using Design Setup.
  • Imported data into Spark RDD from several sources such as HDFS/HBase and built a data pipeline using Kafka to store data in HDFS. Ongoing data was subjected to real-time analysis.
  • Experience building and deploying sophisticated applications and distributed systems in Azure cloud architecture.
  • Involved in converting Hive or SQL queries into Spark transformations using Python and Scala
  • To ingest data from external sources and perform custom processes, I worked on connecting Apache Kafka with the Spark Streaming process.
  • Created the PySpark programs to load the data into Hive and MongoDB databases from PySpark Data frames.
  • Using Spark Context, PySpark, Spark-SQL, Data Frame, and Pair RDD's to investigate and improve the efficiency and optimization of current Hadoop methods.
  • In Scala, I used the Data Frame API to turn a distributed collection of data into named columns.
  • To summarize and alter data, I created Spark jobs and Hive jobs.
  • Using partitions, broadcasts in Spark, effective and efficient joins, and transformations throughout the ingestion process, performance improvement while dealing with huge datasets is possible.
  • Used spark for interactive searches, streaming data processing, and data integration with NoSQL databases such as HBase.
  • Used Sqoop to extract and load data into a Data Lake environment (MS Azure), which was then accessible by business users.
  • Using Hive tables and Hive Serdes, I stored the data in tabular representations.
  • Partitioning, Dynamic Partitions, and Bucketing were implemented in Hive for efficient data access.
  • HBase tables were redesigned to increase speed based on query needs.
  • Map that has been developed to Reduce the number of Java tasks required to convert data files to the Parquet file format.
  • To the analysts, I developed Hive queries for data sampling and analysis.
  • Executed Hive queries that aided in trend research by comparing fresh data to data warehouse reference tables and previous data.
  • Experience designing and architecting several data pipelines, including end-to-end ETL and ELT processes for data intake and transformation in Azure, as well as coordinating tasks among team members.

Environment: CDH5, Hue, Eclipse, Centos Linux, HDFS, MapReduce, Kafka, Azure cloud platforms (hdinsight, datalake, databricks, Blob Storage, Data Factory, Synapse, SQL, SQL DB, DWH and Data Storage Explorer), AWS, Python, Scala, Java, Hive, Sqoop, Spark, PySpark, SQL, Spark-Streaming, HBase, Oracle10g, Oozie, Red Hat Linux.

Hadoop Developer

Confidential

Responsibilities:

  • Worked on developing ETL pipelines to transfer legacy systems data to a Hadoop cluster.
  • Involved in data cleansing, processing, validation of business requirements, and designing functional specifications for schema and table creations and optimized the query performance using HIVE.
  • Built, configured, and maintained Hadoop cluster using Hortonworks distribution based on business requirements.
  • Leveraged Apache bigdata Hadoop components like HDFS, MapReduce, Yarn, HBase, Sqoop, Pig.
  • Implemented end-to-end ETL pipelines using Python and SQL for high-volume analytics. Reviewed use cases before onboarding to HDFS.
  • Developed application to clean semi-structured data like JSON into structured files before ingesting them into HDFS
  • Proficient in developing big data processing applications using Spark APIs (RDDs, Spark SQL, Data Frames, Datasets, Spark Streaming) using Python and Scala.
  • Created inner instrument for contrasting the RDBMS and Hadoop with the end goal that all the data in source and target matches using shell script and can decrease the complexity in moving data.
  • Responsible for loading, managing, and reviewing terabytes of log files using Ambari web UI.
  • Migrated data from traditional RDBMS to HDFS using Sqoop. Ingested data, from MS SQL, Teradata, and Cassandra databases.
  • Identified required tables, views and exported them to Hive and performed ad-hoc queries using Hive joins, Partitioning, Bucketing techniques for faster data access.
  • Worked on Hive integration with Spark SQL scripts for performance enhancement. Implemented file compression techniques using Spark.
  • Automated data flow between various systems using Nifi, designed dataflow models and complicated target tables to obtain relevant metrics from various sources.
  • Used Kafka and merged the web log data from multiple servers and made it available to downstream systems for data analysis and engineering purposes, Boosted Kafka's performance by implementing security.
  • Experience configuring Zookeeper for clusters to coordinate the servers in a process, and to maintain consistency in the data used to make decisions
  • Developed data pipelines from on premises to cloud using AWS services like DMS, S3, EC2 and Redshift.
  • Configured and built AWS Lambda functions to automate data ingestion and terminate resources not in use.
  • Developed Bash scripts to get log files from the FTP server and executed Hive jobs to parse them.
  • Worked on building continuous integration tools using Jenkins.
  • Automated the Jenkins pipeline’s integration and delivery service using Groovy scripts. Used Jenkins for CI/CD and SVN for version control.
  • Developed reporting dashboards using PowerBI by integrating them to MS SQL Server as the back-end database.

Environment: Spark, Spark-Streaming, Spark SQL, EMR, MapReduce, HDFS, Hive, Pig, Pyspark, Shell scripting, Linux, MySQL, Oracle, Kafka, Python, SQL, Java, Teradata, Oracle, MySQL, Tableau, SVN, JIRA

ETL Developer

Confidential

Responsibilities:

  • Used TSQL, PostgreSQL and MySQL for creating Procedures, Triggers, Indexes, Views, and Packages.
  • Designed selection criteria document and involved in target mappings
  • Worked on design and development of Informatica mappings, workflows to load data into staging area, data warehouse and data marts in SQL Server and Oracle.
  • Using new Data Stage features and Unix Shell scripts, the ETL process was optimized and automated.
  • Writing Impala SQL and Redshift SQL to extract data from HDFS and S3 buckets to outbound data marts is a plus. Hands-on experience with S3 bucket creation and deletion for a variety of applications.
  • In Informatica Power Center, I created mappings, sessions, and workflows to populate data into dimension, fact, and lookup tables from many source systems at the same time (SQL server, Oracle, Flat files).
  • Teradata procedures were written to load incremental/aggregated data from Teradata's Core to Semantic layers.
  • Used Teradata utilities FAST LOAD, MULTI LOAD, TPUMP to load data.
  • To automate procedures and populate parameter files, I used PMCMD commands on the command prompt and Unix Shell scripts.
  • Experience in writing and testing Teradata Fast load, Multi load and BTEQ scripts, DML and DDL commands.
  • Created mappings using various Transformations like Source Qualifier, Aggregator, Expression, Filter, Router, Joiner, Stored Procedure, Lookup, Update Strategy, Sequence Generator and Normalizer.
  • Deployed mapplets to avoid duplicate metadata while reducing the development time.
  • Dimensions have been updated utilizing version mapping, which is slowly changing in order to maintain the history of the targeted database.
  • Involved in migration of Informatica from 8.x to 9.x.
  • Created and Monitored Workflows using Workflow Builder and Monitor.
  • By optimizing source and target bottlenecks and implementing pipeline partitioning, the performance of mapping and sessions were improved.
  • Developed and Partitioned indexes on tables while working with DBA.
  • Involved in Query tuning and interpretation of explain plans to improve performance.
  • Involved in exporting database, table spaces, tables using Data pump as well as traditional export/import.
  • Scheduled batch jobs by writing shell scripts.
  • Involved in developing unit test cases and documenting of the ETLprocess

Environment: Informatica Power Center 9.x/8.x, SQL Server, Teradata, PL/ SQL, Oracle, Toad, Cognos 8.3, Windows NT, UNIX Shell Scripting.

We'd love your feedback!