We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

3.00/5 (Submit Your Rating)

SUMMARY

  • 10 years of experience in IT industry Software Design, Analysis, Data Engineer, Big Data technologies, Development, Testing and Database Development by following standard SDLC like Agile.
  • Experience in building ETL solutions on a defined architecture based on cloud and Big Data applications.
  • S3 as Data Lake which responsible for maintaining and handling data inbound and outbound requests through big data platform.
  • Hands - on experience in developing and implementing Big Data solutions and data mining applications on Hadoop using HDFS, MapReduce, HBase, Pig, Hive and Sqoop, Flume, Kafka, Storm, Spark, Oozie, Zookeeper, Flink, NiFi.
  • Experienced in converting batch data into csv files and stored them into AWS S3 in parquet format by using AWS EC2, then structured and stored in AWS Redshift.
  • Created Azure SQL database, performed monitoring and restoring of Azure SQL database. Performed migration of Microsoft SQL server to Azure SQL database.
  • Experience in delivering and supporting large scale data-management and migration projects.
  • Strong foundation in Data Engineering, Software testing, Database’s and ETL.
  • Experience in Junit, REST Services, Web Services.
  • Experience on using Agile methodology (SCRUM).
  • Experience on Big Data technology like Hadoop, Scala, MapR and Spark.
  • Familiar with git and RTC for source control.

PROFESSIONAL EXPERIENCE

Confidential

Sr. Data Engineer

Responsibilities:

  • Work on AWS Data pipeline to configure data loads from S3 to into Redshift.
  • Extracted, transformed and loaded data from various heterogeneous data sources and destinations using AWS Redshift.
  • Migrate data from on-premises to AWS storage buckets
  • Developed a python script to hit REST API's and extract data to AWS S3
  • Created Tables, Stored Procedures, and extracted data using T-SQL for business users whenever required.
  • Performs data analysis and design, and creates and maintains large, complex logical and physical data models, and metadata repositories using ERWIN and MB MDR
  • Selected and generated data into csv files and stored them into AWS S3 by using AWS EC2 and then structured and stored in AWS Redshift.
  • Designed the schema, configured and deployed AWS Redshift for optimal storage and fast retrieval of data and used Spark Data frames, Spark-SQL, Spark MLLib extensively and developing.
  • Generates ETL scripts to transform, flatten, and enrich the data from source to target using AWS Glue and created event driven ETL pipelines with AWS Glue.
  • Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.
  • Worked on developing Pyspark script to encrypting the raw data by using hashing algorithms concepts on client specified columns.
  • Used PySpark and Pandas to calculate the moving average and RSI score of the stocks and generated them into data warehouse.
  • Developed Spark Applications by using Scala and Implemented Apache Spark data processing Project to handle data from various RDBMS and Streaming sources.
  • Exploring with Spark to improve the performance and optimization of the existing algorithms in Hadoop using Spark context, Spark-SQL, PostgreSQL, Data Frame, OpenShift, Talend, pair RDD's
  • Develop Matillion pipelines to ingest data from cloud applications like salesforce, Azure, AWS to snowflake.
  • Design SNS notifications for Matillion pipelines which will notify on specific success/ failures of crucial data processing states.
  • Implemented AWS SQS queue to create the dependency between Matillion jobs across projects
  • Involved in integration of Hadoop cluster with spark engine to perform BATCH and GRAPHX operations.
  • Performed data preprocessing and feature engineering for further predictive analytics using Python Pandas.
  • Responsible for Design, Development, and testing of the database and Developed Stored Procedures, Views and Triggers.
  • Set up Data Lake in Google cloud using Google cloud storage, BigQuery and Big Table.
  • Coordinated with team and Developed framework to generate Daily adhoc reports and Extracts from enterprise data from BigQuery.
  • Compiling and validating data from all departments and Presenting to Director Operation.
  • Created Tableau reports with complex calculations and worked on Ad-hoc reporting using Tableau.
  • Worked on the tuning of SQL Queries to bring down run time by working on Indexes and Execution Plan.
  • Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
  • Pre-processing using Hive and Pig.
  • Developed a detailed project plan and helped manage the data conversion migration from the legacy system to the target snowflake database.
  • Design, develop, and test dimensional data models using Star and Snowflake schema methodologies under the Kimball method.
  • Developed data pipeline using Spark, Hive, Pig, python, Impala, and HBase to ingest customer
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Strong understanding of AWS components such as EC2 and S3
  • Worked on Ingesting data by going through cleansing and transformations and leveraging AWS Lambda, AWS Glue and Step Functions.
  • Worked on a python script to extract data from Netezza databases and transfer it to AWS S3
  • Used ETL to implement the Slowly Changing Transformation, to maintain Historically Data in Data warehouse.
  • Created a Lambda Deployment function, and configured it to receive events from S3 buckets
  • Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
  • Ensure deliverables (Daily, Weekly & Monthly MIS Reports) are prepared to satisfy the project requirements cost and schedule
  • Experience in deploying code through Jenkins and creating pull requests using bit bucket.
  • Used Git for version control with colleagues.

Confidential - Mooresville, NC

Data Engineer

Responsibilities:

  • Worked on Spark Architecture including spark core, spark SQL, DataFrame, Driver Node, Worker Node, Stages, Executors and Tasks, Deployment modes, the Execution hierarchy, fault tolerance, and collection.
  • Developed ETL Processes in Databricks to extract data from redshift, perform transformations and load data to S3- Datalayer in Databricks.
  • Implemented Unit test cases in Pytest, Acceptance Testing in Gauge Framework using python
  • Good Understanding of Data ingestion, Airflow Operators for Data Orchestration, and other related python libraries.
  • Supported AWS Cloud environment with 2000 plus AWS instances configured Elastic IP and Elastic storage deployed in multiple Availability Zones for high availability.
  • Designed appropriate partitioning/bucketing schema to allow faster data retrieval during analysis.
  • Demonstrated good communication skills and story narratives while Sprint Demos to leadership and Stake holders.
  • Experience in debugging Jenkins’s pipeline for log errors
  • Developed ETL Processes in AWS Glue to migrate data from external sources like S3, ORC/Parquet/ Text Files into AWS Redshift.
  • Develop Matillion pipelines to ingest data from various on - prem databases to snowflake.
  • Troubleshooted and maintained ETL/ELT jobs running using Matillion.
  • Develop Python and SQL used in the transformation process in Matillion.
  • Developed scripts in BigQuery and connecting it to reporting tools.
  • Work related to downloading BigQuery data into Spark data frames for advanced ETL capabilities.
  • Worked on Ingesting data by going through cleansing and transformations and leveraging AWS Lambda, AWS Glue and Step Functions.
  • Implemented usage of Amazon EMR for processing Big Data across a Hadoop Cluster of virtual servers on Amazon Elastic Compute Cloud (EC2) and Amazon Simple Storage Service (S3)
  • Created AWS Lambda functions and assigned IAM roles to schedule python scripts using CloudWatch Triggers to support the infrastructure needs (SQS, Event Bridge, SNS)
  • Developed a python script to hit REST API’s and extract data to AWS S3
  • Conducted ETL Data Integration, Cleansing, and Transformations using AWS glue Spark script
  • Worked on functions in Lambda that aggregates the data from incoming events, and then stored result data in Amazon DynamoDB
  • Worked on No-SQL database like DynamoDB and MongoDB
  • Deployed the project on Amazon EMR with S3 connectivity for setting a backup storage
  • Worked on AWS Data Pipeline to configure data loads from S3 to into Redshift
  • Used JSON schema to define table and column mapping from S3 data to Redshift.
  • Connected Redshift to Tableau for creating dynamic dashboard for analytics team.
  • Used JIRA to track issues and Change Management.
  • Involved in creating Jenkins jobs for CI/CD using git, Maven and Bash scripting.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs with Scala.
  • Used Spark API over Hadoop YARN as execution engine for data analytics using Hive.
  • Exported the analyzed data to the relational databases using Sqoop to further visualize and generate reports for the BI team.
  • Migrated the computational code in hql to PySpark.
  • Worked with Spark Ecosystem using Scala and Hive Queries on different data formats like Text file and parquet.
  • Worked in migrating Hive QL into Impala to minimize query response time.
  • Responsible for migrating the code base to Amazon EMR and evaluated Amazon eco systems components like Redshift.
  • Collected the logs data from web servers and integrated in to HDFS using Flume
  • Developed Python scripts to clean the raw data.
  • Imported data from AWS S3 into Spark RDD, Performed transformations and actions on RDD's
  • Used AWS services like EC2 and S3 for small data sets processing and storage
  • Implemented Nifi flow topologies to perform cleansing operations before moving data into HDFS.
  • Worked on different file formats (ORCFILE, Parquet, Avro) and different Compression Codecs (GZIP, SNAPPY, LZO).
  • Created applications using Kafka, which monitors consumer lag within Apache Kafka clusters.
  • Worked on importing and exporting data into HDFS and Hive using Sqoop, built analytics on Hive tables using Hive Context in spark Jobs.
  • Developed workflow in Oozie to automate the tasks of loading the data into HDFS.
  • Worked in Agile environment using Scrum methodology.
  • Work on requirements gathering, analysis and designing of the systems.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster testing and processing of data.
  • Involved in designing Kafka for multi data center cluster and monitoring it.
  • Responsible for importing real time data to pull the data from sources to Kafka clusters.
  • Worked with spark techniques like refreshing the table and handling parallelly and modifying the spark defaults for performance tuning.
  • Implemented Spark RDD transformations to Map business analysis and apply actions on top of transformations.
  • Involved in migrating MapReduce jobs into Spark jobs and used SparkSQL and Data frames API to load structured data into Spark clusters.
  • Involved in using Spark API over Hadoop YARN as execution engine for data analytics using Hive and submitted the data to BI team for generating reports, after the processing and analyzing of data in Spark SQL.
  • Performed SQL Joins among Hive tables to get input for Spark batch process.
  • Worked with data science team to build statistical model with Spark MLLIB and PySpark.
  • Involved in performing importing data from various sources to the Cassandra cluster using Sqoop.
  • Worked on creating data models for Cassandra from Existing Oracle data model.
  • Used Sqoop to import functionality for loading Historical data present in RDBMS to HDFS.
  • Designed Column families in Cassandra and Ingested data from RDBMS, performed data transformations, and then export the transformed data to Cassandra as per the business requirement.
  • Configured Hive bolts and written data to hive in Hortonworks as a part of POC.
  • Implemented ELK (Elastic Search, Log stash, Kibana) stack to collect and analyze the logs produced by the spark cluster.
  • Developed Python script for start a job and end a job smoothly for a UC4 workflow.
  • Developed Oozie workflow for scheduling & orchestrating the ETL process.

Confidential - Langhorne, PA

Data Engineer

Responsibilities:

  • Expertise in High Availability and Recoverability of databases using standard MS SQL technologies including
  • Design and implement database solutions in Azure SQL Data Warehouse, Azure SQL.
  • Design & implement migration strategies for traditional systems on Azure (Lift and shift/Azure Migrate, other third-party tools.
  • Design Setup maintain Administrator the Azure SQL Database, Azure Analysis Service, Azure SQL Data warehouse, Azure Data Factory, Azure SQL Data warehouse.
  • Used Azure Data Factory extensively for ingesting data from disparate source systems.
  • Used Azure Data Factory as an orchestration tool for integrating data from upstream to downstream systems.
  • Automated jobs using different triggers (Event, Scheduled and Tumbling) in ADF.
  • Migrate SQL Server and Oracle database to Microsoft Azure Cloud.
  • Migrate the Data using Azure database Migration Service (AMS).
  • Worked Azure SQL Database Environment.
  • Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform and load data from different sources like Azure SQL, Azure SQL Data warehouse, write-back tool and backwards.
  • Manage SQL Server databases through multiple project lifecycle environments, from development to mission-critical production systems.
  • Have good experience working with Azure Data lake storage and loading data into Azure SQL Synapse analytics (DW).
  • Involved Cosmos DB for storing catalog data and for event sourcing in order processing pipelines.
  • Designed and developed user defined functions, stored procedures, triggers for Cosmos DB
  • Analyzed the data flow from different sources, target to provide the corresponding design Architecture in Azure environment.
  • Created Application Interface Document for the downstream to create new interface to transfer and receive the files through Azure Data Share.
  • Creating pipelines, data flows and complex data transformations and manipulations using ADF and PySpark with Databricks
  • Ingested data in mini-batches and performs RDD (Resilient Distributed Dataset) transformations on those mini-batches of data by using Spark Streaming to perform streaming analytics in Data bricks.
  • Created, provisioned different Databricks clusters needed for batch and continuous streaming data processing and installed the required libraries for the clusters.
  • Integrated Azure Active Directory authentication to every SQL request sent and demoed feature to Stakeholders
  • Improved performance by optimizing computing time to process the streaming data and saved cost to company by optimizing the cluster run time.
  • Designed and developed a new solution to process the NRT data by using Azure stream analytics, Azure Event Hub and Service Bus Queue.
  • Created Linked service to land the data from SFTP location to Azure Data Lake.
  • Created numerous pipelines in Azure using Azure Data Factory v2 to get the data from disparate source systems by using different Azure Activities like Move &Transform, Copy, filter, for each, Databricks etc.
  • Good experience in creating Elastic pool databases and schedule Elastic jobs for executing TSQL procedures.
  • Worked on creating tabular models on Azure analysis services for meeting business reporting requirements.

Confidential

Data Engineer

Responsibilities:

  • Working with complex SQL, Stored Procedures, Triggers, and packages in large databases from various servers.
  • Involved in complete Software Development Life Cycle (SDLC) process by analyzing business requirements and understanding the functional workflow of information from source systems to destination systems.
  • Import data using Sqoop to load data from Teradata to HDFS on a regular basis.
  • Write Hive queries for ad-hoc reporting to the business.
  • Develop and Orchestrated Hadoop jobs using Oozie job scheduler.
  • Develop Streaming applications using pyspark for building stream data platform integrating with Kafka.
  • Worked on developing pyspark applications to ingest streaming data from Kafka topics into HDFS Data lakes.
  • Develop pyspark applications to apply business validation rules to incoming transactional data.
  • Develop data processing applications using pyspark to process transactional data and persist in data lakes.
  • Develop HQL Scripts to create external tables in Hive on top of ingested data and processed data.
  • Develop pyspark applications to join data transactional data with multiple dimensional tables, process the data, and persist the data to Cassandra.
  • Create analytical applications to perform certain analytics and push the data RDBMS for business analysis.
  • Built an Ingestion Framework that would ingest the files from SFTP to HDFS using Apache nifi.
  • Expert in writing, configuring and maintaining the Hibernate configuration files and writing and updating Hibernate mapping files for each Java object to be persisted.
  • Involved in application performance tuning and fixing bugs.
  • Performed various pocs in data ingestion, data analysis, and reporting using Hadoop, MapReduce, Hive, Pig, Sqoop, Flume, Elastic Search.
  • Research and recommend various tools and technologies on the Hadoop stack considering the workloads of the organization.
  • Extensively used SQL queries, PL/SQL stored procedures & triggers in data retrieval and updating of information in the Oracle database using JDBC.
  • Enhanced and optimized the functionality of Web UI using AJAX, XSL, XSLT, CSS, XHTML, and JavaScript.
  • Developed Web Services using JAX-RPC, JAXP, WSDL, SOAP, XML to provide the facility to obtain quotes, receive updates to the quote, customer information, status updates, and confirmations.
  • Analyzing the business requirements and doing the GAP analysis then transforming them into detailed design specifications.
  • Performed Code Reviews and responsible for Design, Code, and Test signoff.
  • Assigning work to the team members and assisting them in development, clarifying on design issues, and fixing the issues.
  • Work with project teams to review reporting / data analysis needs.
  • Create ETL data pipelines using Azure data factory to load data into Azure SQL Server data warehouse.
  • Create Azure SQL Server data warehouse objects; tables, views, and stored procedures as needed to support data analytics processes.
  • Create COSMOS/SCOPE scripts to extract data from big data streams for analysis.
  • Create SQL scripts to support Powerbi reports and adhoc data requests.
  • Create Power BI dashboard reports utilizing DAX and interactive report visualizations.
  • Performed Deep dive data using MS Excel tools, Pivot tables, Lookups, and charts.
  • Migrate legacy systems such as Xflow/cosmos jobs into Azure environment.
  • Perform other processes needed to maintain the data analytic platform such as
  • Azure resource management, DevOPs repositories maintenance, JSON, XML, XFlow scheduled jobs maintenance.

Confidential

Data Analyst

Responsibilities:

  • Devised simple and complex SQL scripts to check and validate Dataflow in various applications.
  • Performed Data Analysis, Data Migration, Data Cleansing, Transformation,
  • Integration, Data Import, and Data Export through Python.
  • Devised PL/SQL Stored Procedures, Functions, Triggers, Views and packages. Made use of Indexing, Aggregation and Materialized views to optimize query performance.
  • Developed logistic regression models (using R programming and Python) to predict subscription response rate based on customers variables like past transactions, response to prior mailings, promotions, demographics, interests and hobbies, etc.
  • Created Tableau dashboards/reports for data visualization, Reporting and Analysis and presented it to Business.
  • Created Data Connections, Published on Tableau Server for usage with Operational or Monitoring Dashboards.
  • Knowledge in Tableau Administration Tool for Configuration, adding users, managing licenses and data connections, scheduling tasks, embedding views by integrating with other platforms.
  • Worked with senior management to plan, define and clarify dashboard goals, objectives and requirement.
  • Responsible for daily communications to management and internal organizations regarding status of all assigned projects and tasks.
  • Responsible to tune ETL procedures and schemas to optimize load and query.

We'd love your feedback!