Sr Data Engineer Resume
Charlotte, NC
SUMMARY
- 8+ years of experience on Bigdata in implementing end - to-end Hadoop solutions and analytics using various Hadoop distributions like Cloudera Distribution of Hadoop (CDH), Hortonworks (HDP), and MapR distribution.
- Experience in Apache Hadoop ecosystem components like Spark, Scala, HDFS, HBase, Kafka, Hive, MapReduce, Sqoop, Zookeeper, Airflow, Snowflake, YARN, Impala, Flume, Pig, Nifi, and Oozie.
- Experience in developing Spark applications using Spark-Core, Spark Context,Spark-SQL,Spark Streaming, Data frames,PairRDD and knowledge of Spark MLLib.
- Experience in AWS cloud services (VPC, EC2, S3, RDS, CLI, Athena, Glue, Redshift, IAM, Data Pipeline, EMR, DynamoDB, Elastic Search, Lambda, Elastic Load Balancing, Auto Scaling, Cloud Front, CloudWatch, Kinesis, SNS, SES, and SQS).
- Experience in various relational databases such as SQL, PostgreSQL, MySQL, Oracle, DB2, MS Access, and NoSQL databases such asMongoDB,HBase,Cassandra,DynamoDB, and Redshift.
- Experience in creating Data frames using PySpark and performing operations on the Data frames using Python
- Experience in building and optimizing big data pipelines, architectures, and data sets like Hadoop, Spark, and Hive.
- Experienced in analyzing data using HIVEQL, PIG Latin and custom MapReduce programs in Java and extending HIVE and PIG core functionality by using custom UDF’s.
- Extensive experience in Data Visualization including producing tables, graphs, listings using various procedures and tools such as Tableau.
- Expertise in Creating, Debugging, Scheduling and Monitoring jobs using Airflow for ETL batch processing to load into Snowflake for analytical processes.
- Experience in different file formats like Avro, Parquet, ORC, XML, JSON.
- Experienced in handling various types of data (Structured, Semi-Structured, Unstructured) coming from various sources with both batch and real time data pipelines.
- Experience in working with CI/CD pipeline using tools like Jenkins, Terraform, Docker, Kubernetes, Ansible, and Chef.
- Experience in various build tools such as Ant, Maven, and Gradle.
- Expertise in Creating, Debugging, Scheduling and Monitoring jobs using Airflow and Oozie.
- Experienced on implementation of a log producer in Scala that watches for application logs, transform incremental log, and sends them to a Kafka and Zookeeper based log collection platform.
- Experience in working parallelly in both GCP and Azure clouds coherently.
- Experience in Dimensional Modeling (OLAP), Data Migration, Data Cleansing, Data Profiling, and ETL Processes features for data warehouses.
- Experienced writing Test cases and implement unit test cases using testing frameworks like Junit, Easy mock and Mockito.
- Hands on experience on Data Analytics Services such as Athena, Glue Data Catalog & Quick Sight.
- Experience in working with various version control tools like SVN, CVS, and Git.
- Experience in project management and Bug Tracking tools such as JIRA and Bugzilla.
- Experience in delivering the complex projects using Agile and Scrum methodology.
TECHNICAL SKILLS
Big Data Ecosystems: Hadoop, Map Reduce, Spark, HDFS, HBase, Pig, Hive, Sqoop, Kafka, Cloudera, Horton works, Oozie, Nifi, and Airflow.
Spark Technologies: Spark SQL, Spark Data frames and RDD
Scripting Languages: Python and shell scripting
Programming Languages: Python, Scala, SQL, PL/SQL
Cloud Technologies: Azure, AWS EMR, EC2, S3, Glue, Athena, Redshift, Docker
Databases: Oracle, MySQL and Microsoft SQL Server
NoSQL Technologies: HBase, MongoDB, Cassandra, DynamoDB
BI tools: Tableau, Kibana, Power BI
Web Technologies: SOAP, and REST.
Other Tools: Eclipse, PyCharm, Git, ANT, Maven, Jenkins, SOAP UI, QC, Jira, Bugzilla, Palantir Foundry
Methodologies: Agile /Scrum, Waterfall
Cloud Technologies: Azure, AWS, GCP
Operating Systems: Windows, UNIX, LINUX.
PROFESSIONAL EXPERIENCE
Confidential, Charlotte, NC
Sr Data Engineer
Responsibilities:
- Performed Data Analysis, Data Migration, Data Cleansing, Transformation, Integration, Data Import, and Data Export throughPython.
- UtilizedSparkSQL API inPySparkto extract and load data and perform SQL queries.
- Worked on developingPysparkscript to encrypting the raw data by using hashing algorithms concepts on client specified columns.
- Exploring with Spark to improve the performance and optimization of the existing algorithms in
- Hadoop using Spark context, Spark-SQL, PostgreSQL, Data Frame, OpenShift, Pair RDD’s.
- Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing.
- Involved in designing and deploying multi-tier applications using all the AWS services like (EC2, Route53, S3, Lambda, Cloud Watch, RDS, Dynamo DB, SNS, SQS, IAM) focusing on high availability, fault tolerance, and auto-scaling in AWS Cloud Formation.
- Worked on Kafka REST API to collect and load the data on Hadoop file system and used sqoop to load the data from relational databases.
- Involved in Installing, configuring, and maintaining Data Pipelines.
- Developed Spark/Scala, Python for regular expression (regex) project in the Hadoop/Hive environment with Linux/Windows for big data resources.
- Developed a NiFi Workflow to pick up the data from Data Lake as well as from server and send that to Kafka broker.
- Built various jobs using AWS Glue’s crawler capabilities to perform data cataloguing and building the ETL pipelines for the target data mart.
- Performed the migration of Hive and MapReduce Jobs from on - premise MapR to AWS cloud using EMR.
- Involved in monitoring servers using Nagios, Cloud watch and using ELK Stack Elastic search Kibana.
- Used Pentaho data transformation steps to transform the data into client DSF (data standard format).
- Created Performance Testing program in Python Spark to compare between two big data result sets.
- Responsible in building the data ingestion pipelines using AWS EMR and Spark Scala as the Data Processing Engine and AWS Athena as the Consumption Layer
- Supporting Continuous storage in AWS using Elastic Block Storage, S3, Glacier, and Created Volumes and configured Snapshots for EC2 instances.
- Involved in building a data pipeline and performed analytics usingAWS stack(EMR, EC2, S3, RDS, Lambda, Kinesis, Athena, ELB, Glue, SQS, Redshift, and ECS).
- Designed, developed, and implemented pipelines using python API (PySpark) of Apache Spark on AWS EMR.
- Used PySpark and Pandas to calculate the moving average and RSI score of the stocks and generated them into data warehouse.
- Created scripts to readCSV, JSON and parquet filesfrom S3 buckets inPythonand load intoAWS S3, DynamoDB (NoSQL) and Snowflake.
- Performed data analysis on the raw data residing on HDFS and local storage using Hive queries and Sqoop to import data between Hadoop and RDBMS.
- Generated report on predictive analytics using Python and Tableau including visualizing model performance and prediction results.
- Involved in developing and documenting the ETL strategy to populate the data warehouse from various source systems.
- Integrated Apache Airflow with AWS to monitor multi-stage ML workflows with the tasks running on Amazon Sage Maker.
- Worked on Dimensional and Relational Data Modelling using Star and Snowflake Schemas, OLTP/OLAP system, Conceptual, Logical and Physical data modelling using Erwin.
- Automated the data processing with Oozie to automate data loading into the Hadoop Distributed File System (HDFS).
- Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, MongoDB, T-SQL, and SQL Server using Python.
- Utilized Agile and Scrum methodology for team and project management.
- Used Jenkins for CI/CD and Git as a version control tool.
Environment: Hadoop, Spark, Scala, Kafka, Python, Nifi, Snowflake, Airflow, Tableau, PostgreSQL, HBase, HDFS, NoSQL (MongoDB), Pentaho, ETL, GitHub, AWS, AWS Glue, Athena, Palantir Foundry, Sage Maker, Pig, Hive, Agile.
Confidential, Boston, MA
Azure Data Engineer
Responsibilities:
- Analyze, design and build Modern data solutions using Azure PaaS service to support visualization of data. Understand current Production state of application and determine the impact of new implementation on existing business processes.
- Worked on analyzing Hadoop cluster using different big data analytic tools including Flume, Pig, Hive, HBase, Oozie, Zookeeper, Sqoop, Spark and Kafka.
- Worked in creating HDInsight cluster and Storage Account with End-to-End environment for running the jobs.
- Played a lead role in the development of Confidential Data Lake and in building Confidential Data Cube on Microsoft Azure HDINSIGHT cluster.
- Worked with ETL tools including Talend Data Integration, Talend Bigdata, Pentaho data integration.
- Involved in handling Python and Spark Context when writing Pyspark programs for ETL.
- Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL and U-SQL Azure Data Lake Analytics . Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in In Azure Databricks.
- Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
- Involved in writing Google dataflow pipelines and transformation in presentation layer.
- Designed and implemented configurable data delivery pipeline for scheduled updates to customer facing data stores built with Python.
- Involved in building the data pipelines in GCP for ETL related jobs.
- Carried out data transformation and cleansing using SQL queries, Python, and Pyspark.
- Developed Spark applications using Pyspark and Spark-SQL for data extraction, transformation and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns.
- Created ETL scripts to reteive data feeds, page metrics from Google Analytic services.
- Migrated previously written cron jobs to airflow/composer in GCP.
- Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
- To meet specific business requirements wrote UDF’s in Scala and Pyspark.
- Developed JSON Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the SQL Activity.
- Implemented OLAP multi-dimensional cube functionality using AzureSQL Data Warehouse.
- Hands-on experience on developing SQL Scripts for automation purpose.
- Created Build and Release for multiple projects (modules) in production environment using Visual Studio Team Services (VSTS).
- Wrote AzurePower shellscripts to copy or move data from local file system to HDFS Blob storage.
- Worked extensively with Dimensional modeling, Data migration, Data cleansing, ETL Processes for data warehouses.
- Worked in Agile Methodology and used JIRA for maintain the stories about project.
- Involved in gathering the requirements, designing, development and testing.
Environment: Hadoop, Azure Data Factory, Azure Data Lake, Azure Storage, Azure SQL, Azure DataWarehouse, Azure Databricks, Azure Power Shell, ETL, GCP, Map Reduce, Palantir Foundry, Hive, Spark, Python, Yarn, Tableau, Kafka, Sqoop, Scala, HBase, HDINSIGHT.
Confidential, Seattle, WA
Data Engineer
Responsibilities:
- Developed data pipelines using Sqoop, Pig and Hive to ingest customer data into HDFS to perform data analytics.
- Familiar with data architecture including data ingestion pipeline design, data modelling and data mining.
- Developed Hive queries to pre-process the data required for running the business process.
- Worked on developing ETL processes to load data from multiple data sources to HDFS using SQOOP.
- ConfiguredSpark streamingto get ongoing information from theKafkaand store the stream information to HDFS.
- Import the data from different sources like HDFS/HBase into Spark RDD and perform computations using PySpark to generate the output response.
- Used Talend for Big data Integration using Spark and Hadoop.
- Developing Spark scripts, UDFS using both Spark DSL and Spark SQL query for data aggregation, querying, and writing data back into RDBMS through Sqoop.
- Written multiple MapReduce Jobs using Java API, Pig for data extraction, transformation and aggregation from multiple file formats including Parquet, Avro, XML, JSON, CSV, ORCFILE.
- Designed and developed Security Framework to provide fine grained access to objects in AWS S3 using AWS Lambda, DynamoDB (NoSQL).
- Developed a detailed project plan and helped manage the data conversion migration from the legacy system to the target snowflake database.
- Used DataStaxSpark connector which is used to store the data into Cassandra database or get the data from Cassandra database.
- Utilized Kubernetes and Docker for the runtime environment of theCI/CDsystem to build, test deploy.
- Involved in loading and transforming large sets of structured data from router location to EDW using an Apache NiFi data pipeline flow.
- Responsible for analysis of requirements and designed generic and standard ETL process to load data from different source systems.
- Performed end- to-end Architecture & implementation assessment of various AWS Cloud services like Amazon EMR, Redshift, S3, IAM, RDS, Cloud Watch, Athena.
- Performing statistical data analysis and data visualization using Python.
- Worked onMongoDB (NoSQL)for distributed storage and processing.
- Used AWS EMR to transform and move large amounts of data into and out of other AWS data stores and databases, such as Amazon Simple Storage Service and Amazon DynamoDB.
- Creating Lambda functions with Boto3 to deregister unused AMIs in all application regions to reduce the cost for EC2 resources.
- Implementing POC to migrate map reduce jobs into Spark RDD transformations using Python.
- Transformed the data using AWS Glue dynamic frames with PySpark; cataloged the transformed the data using Crawlers and scheduled the job and crawler using workflow feature.
- Developed reusable framework to be leveraged for future migrations that automates ETL from
- RDBMS systems to the Data Lake utilizing Spark Data Sources and Hive data objects.
- Conducted Data blending, Data preparation using Alteryx and SQL for Tableau consumption and publishing data sources to Tableau server.
- Worked on scheduling all jobs using Airflow scripts using python added different tasks to DAG, LAMBDA.
- Developed Kibana Dashboards based on the Log stash data and Integrated different source and target systems into Elastic search for near real time log analysis of monitoring End to End transactions.
- Deploy new hardware and software environments required for PostgreSQL/Hadoop and expand existing environment.
- Implemented AWS Step Functions to automate and orchestrate the Amazon Sage Maker related tasks such as publishing data to S3, training ML model and deploying it for prediction.
- Used Jenkins for CI/CD, Docker as a container tool and Git as a version control tool.
- Followed agile methodology including, test-driven and pair-programming concept.
Environment: Hadoop, Spark, Scala, AWS EMR, S3, RDS, Redshift, Lambda, Boto3, DynamoDB (NoSQL), Sage Maker, Glue, Athena, HBase, Palantir Foundry, ETL, HDFS, Kafka, HIVE, SQOOP, Map Reduce, Pig, Python, Agile, Tableau.
Confidential
Software Engineer
Responsibilities:
- Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
- Analysed the SQL scripts and designed the solution to implement using Scala.
- Used Spark-SQL to Load JSON data and create Schema RDD and loaded it into Hive Tables and handled structured data using Spark SQL.
- Involved in converting MapReduce programs into Spark transformations using Spark RDDs using Scala and Python.
- Involved in designing different components of the system like Sqoop, Hadoop process involves Mapreduce & Hive, Spark, FTP integration to down systems.
- Involved in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for analysing & transforming the data to uncover insights into the customer usage patterns.
- Developed ETLs using PySpark and used both Data frame API and Spark SQL API.
- Using Spark, performed various transformations and actions and the result data is saved back to HDFS from there to the target database Snowflake.
- Processed the Web server logs by developing Multi-hop flume agents by using Avro Sink and loaded them into MongoDB (NoSQL) for further analysis, also extracted files from MongoDB through Flume and processed them.
- Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR
- Involved in real-time data analytics using Spark Streaming, Kafka, and Flume
- Configured Spark streaming to get ongoing information from Kafka and store the stream information to HDFS.
- Installed and configured pig, written Pig Latin scripts to convert the data from a Text file to Avro format.
- Created Partitioned Hive tables and worked on them using HiveQL.
- Design and Develop ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
- Create, modify and execute DDL in table AWS Redshift and snowflake tables to load data.
- Worked in building ETL pipeline for data ingestion, data transformation, and data validation on cloud service AWS, working along with data steward under data compliance.
- Used Pyspark for extracting, filtering, and transforming the Data in data pipelines.
- Involved in monitoring servers using Nagios, Cloud watch, and using ELK Stack Elastic search Kibana
- Used Data Build Tool for transformations in the ETL process, AWS lambda, and AWS SQS.
- Worked on scheduling all jobs using Airflow scripts using python. Adding different tasks to DAG’s and dependencies between the tasks.
- Worked on various automation tools like GIT, Terraform, Docker, and Ansible.
- Created Unix Shell scripts to automate the data load processes to the target Data Warehouse.
- Used Jira for ticketing and tracking issues and Jenkins for continuous integration (CI) and continuous deployment (CD).
- Estimates and planning of development work using Agile Software Development.
Environment: Hadoop, Spark, Scala, AWS, Redshift, Hadoop, Map Reduce, HDFS, Hive, Shell ScriptSQOOP, Python, PostgreSQL, ETL, MongoDB (NoSQL), Airflow, Snowflake, Agile.
Confidential
Data Engineer
Responsibilities:
- Developed business logic using Kafka Direct Stream in Spark Streaming and implemented business transformations.
- Orchestrate data workflows using airflow to manage and schedule by creating DAGS using Python.
- Involved in building PySpark and Spark-Scala applications for interactive analysis, batch processing, and stream processing
- Worked on the Spark SQL and Spark Streaming modules of Spark and used Scala to write code for all Spark use cases.
- Used Sqoop to import data into HDFS and Hive from the Oracle database.
- Pulling the data from the data lake (HDFS) and massaging the data with various RDD transformations.
- Used AWS Cloud with Infrastructure Provisioning / Configuration. - -
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting on the dashboard.
- Maintained Tableau functional reports based on user requirements.
- Working with Spark various modules of Spark Data Frames, RDD, and Spark Context.
- Worked onApache NiFilike executingSpark script, and Sqoop scripts throughNiFi, and worked on creating scatter and gather patterns inNiFi.
- Supporting Continuous storage in AWS using Elastic Block Storage, S3, Glacier, Created Volumes, and configured Snapshots for EC2 instances.
- Managing the various ETL development, and release management activities.
- Worked on providing user support and application support on Hadoop/Big Data infrastructure.
- Worked on evaluating, and comparing different tools for test data management with Hadoop.
- Involved in creating Hive tables, loading with data, and writing hive queries which will run internally in the map-reduced way.
- Worked extensively with Dimensional modeling, Data migration, Data cleansing, and ETL Processes for data warehouses.
- Building data visualizations to summarize the conclusion of advanced analysis.
- BuildingDataPipeline by sourcing and collatingdatafrom different sources and APIs for data consumption.
- Designed and developed ETL mapping for data collection from various data feeds using Rest API.
- Developed storytelling dashboards in Tableau Desktop and published them on to Tableau Server which allowed end users to understand the data.
- Involved inS3 event notifications, an SNS topic, an SQS queue, and a Lambda function sending a message to the Slack channel.
- Transformed Teradata scripts and stored procedures to SQL and Python running on Snowflake’s cloud platform.
- Worked on the agile environment, used GitHub for version control and Team city for the continuous build.
- Continuously monitoring and testing the system to ensure optimized performance.
- Performed CI/CD operations with Gitlab pipelines, Jenkins, Docker, and Kubernetes.
Environment: Hadoop, Spark, Hive, Pig, Spark SQL, Python, Spark Streaming, HBase, NoSQL, HDFS, Sqoop, Kafka, AWS EC2, S3, ETL, Scala, Linux Shell Scripting, Agile
