We provide IT Staff Augmentation Services!

Sr Data Engineer Resume

0/5 (Submit Your Rating)

D, C

SUMMARY

  • Around 7+ years of professional software experience with expertise in Big Data technologies, Hadoop Ecosystem, Cloud Engineering, Data Warehousing, Data Pipelines, SQL/NoSQL, Cloud based RDS, Distributed Database, Serverless Architecture, Data Mining, and cloud technologies like AWS EMR, Redshift, Lambda, Step Functions, Cloud Watch.
  • Expertise in designing, maintaining various applications utilizing Amazon Web Services like AWS Glue, DynamoDB, Kinesis, EC2, S3, EBS, ELB, CloudWatch, RDS, SNS, SQS, IAM, Red Shift, CloudTrail, CloudFormation Template focusing on high availability, fault - tolerance and auto scaling.
  • Proven expertise in deploying major software solutions for clients meeting the business requirements such as Big Data Processing, Ingestion, Analytics and Cloud Migration from On-premises to AWS Cloud.
  • Deployed AWS Kinesis based consumers in Lambda and pipeline data to a data lake while allowing real time analytics using DynamoDB.
  • Worked with AWS BOTO3 to write Python scripts for encrypting EBS volumes, backups and scheduling Lambda functions for running the servers, starting and stopping of EC2 instances and taking snapshots of the servers.
  • Expertise in writing AWS CloudFormation templates in JSON to use them as blueprints for building & deploying multiple AWS resources.
  • Responsible to migrate legacy applications to AWS & Azure clouds as well as migration to SaaS solutions.
  • Work experience in Azure App & Cloud Services, PaaS, Azure Blob Storage, VM creation, ARM Templates, PowerShell scripts, IaaS, storage, network and database.
  • Hands-on experience on Azure Cloud services (Paas), Management tools, Migration, Storage, Network &Content Delivery, Active Directory, Azure container service, VPN Gateway, Content Delivery Management, Azure Storage Services, Azure Database Services.
  • Extensively used Azure services ADLS for storing data and ADF pipelines for resource- intensive jobs triggering.
  • Good experience in building pipelines using Azure Data Factory and moving the data into Azure Data Lake Store
  • Extracting, Parsing, Cleaning, and ingesting the incoming web feed data and server logs into the HDInsight and Azure Data Lake Store by handling structured and unstructured data.
  • Experience in working with Cloudera and Hortonworks Hadoop distros and, to fully leverage and implement new Hadoop features.
  • Good knowledge in Kinesis Data Streams & Kinesis Firehouse & integration with AWS Lambda for serverless data collection.
  • Developing data processing tasks using PySpark such as reading data from external sources, merge data, perform data enrichment and load in to target data destinations.
  • Experienced with Spark Streaming, SparkSQL Kafka & AWS Kinesis for real - time data processing.
  • Highly proficient in writing lambda functions to automate tasks on AWS using CloudWatch triggers, S3 events as well as DynamoDB streams and kinesis streams.
  • Created Airflow DAGs for Batch Processing to orchestrate Python data pipelines for csv files preparation pre-ingestion.
  • Migrating Hive & MapReduce jobs to EMR with automating the workflows using Airflow.
  • Experience building and optimizing ‘big data’ data pipelines, architectures and data sets. (HiveMQ, Kafka, Cassandra, S3, Redshift)
  • Used Kafka and Kafka brokers, initiated the spark context and processed live streaming information with RDD and Used Kafka to load data into HDFS and NoSQL databases.
  • Worked on deploying SQL database to Virtual Machines involving Azure Tables for non-relational data.
  • Experience with NoSQL databases such as Hbase, DynamoDB and Mongo DB.
  • Cleaned and stacked data from client using Python and SQL to build marketing models that resulted in improving on investment.
  • Hands-on experience in implementing and deploying (Elastic Map Reduce) EMR cluster leveraging Amazon Web Services with EC2 instances.
  • Experienced with version control systems like GIT and used Source code management client tools like GitHub and GitLab.
  • Experienced on Hadoop/Hive on AWS, using both EMR and non-EMR-Hadoop in EC2.
  • Hands on experience on Data Analytics Services such as Athena, Glue Data Catalog & Quick Sight.
  • Worked on ETL Migration services by developing and deploying AWS Lambda functions for generating a serverless data pipeline which can be written to Glue Catalog and can be queried from Athena. Ability to develop spark application using PySpark API to analyze huge datasets.
  • Good knowledge of Data Warehouse/Data Mart, OLTP, OLAP, Data Modeling

TECHNICAL SKILLS

Hadoop/BigData Technologies: Hadoop, MapReduce, Oozie, Sqoop, Hive, Spark, and Kafka, Cloudera Manager, Flume.

NO SQL Database: Dynamo DB, Mongo DB, HBase

ETL Tools: Elluminate, Mapper

AWS Cloud: Athena, Kinesis, Glue, EC2, S3, EBS, ELB, Cloud-Front, IAM, RDS, EMR, Elasticsearch, Cloud Watch

Azure Cloud: Azure Data Bricks, HDInsight, ADW and hive

Hadoop Distribution: Cloudera, Horton Works

Scripting Languages: Python, Scala, SQL, PySpark

Version Control: GIT

Databases: Teradata, Oracle, MY SQL

Cluster Managers: Kubernetes, Docker

Operating Systems: Unix, Linux, Mac OS-X, Windows 10, Windows 7, CentOS

Reporting: Tableau, PowerBI, Quick Insight

Development Methodologies: Waterfall, Agile

PROFESSIONAL EXPERIENCE

Confidential, D.C

Sr Data Engineer

Responsibilities:

  • Expertise in Amazon Web Services (AWS) Cloud Platform which includes services like Kinesis, Athena, DynamoDB, Cloud Front, EC2, S3, VPC, ELB, CloudWatch,Security Groups, Red shift, CloudFormation.
  • Built AWS server for deployment and data processing. Involved in entire lifecycle of the projects including Design and Deployment, Testing and Implementation and support. Maintaining the scripts using the version Control.
  • Created Airflow DAGs for Batch Processing to orchestrate Python data pipelines for csv files preparation pre-ingestion, using configuration files to parameterize for a multitude of input files and separate Task Instances.
  • Worked on ETL Migration services by developing and deploying AWS Lambda functions for generating a serverless data pipeline which can be written to Glue Catalog and can be queried from Athena.
  • Created database tables that can store and retrieve any amount of data, and serve any level of request traffic using DynamoDB and Proficiency in multiple databases like MySQL, ORACLE & MS SQL Server.
  • Wrote Python scripts to migrate data into from multiple sources to AWS Redshift.
  • Wrote Python script using AWS S3 Boto3 library to download and upload files from AWS S3 buckets.
  • Deployed Python script in SSIS package for ETL processing through SQL stored procedures
  • Created on-demand tables on S3 files using Lambda Functions and AWS Glue using Python and PySpark.
  • Wrote Lambda functions in python for AWS Lambda and deployed python scripts for data transformations and analytics on large data sets in EMR clusters and AWS Kinesis data streams.
  • Created new connections through applications for better access to MySQL database and involved in writing SQL & PLSQL - Stored procedures, functions, sequences.
  • Managed datasets using Panda data frames and MySQL, queried MYSQL database queries from python using Python-MySQL connector and MySQL DB package to retrieve information.
  • Wrote a multithreaded process to migrate records from database to AWS S3 and Dynamo DB tables for state storage of activities.
  • Created tables on AWS S3 on top of data obtained from different datasources and schedule it using Airflow. Implemented Airflow to manage PySpark data pipeline dependencies utilizing DAG, PythonOperator, HiveOperator and Scheduler
  • Worked on deployment, data security and troubleshooting of the applications using AWS services Analyzed SQL scripts and designed the solutions to implement using PySpark.
  • Wrote conversion scripts using SQL, PL/SQL, stored procedures, functions and packages to migrate data from SQL server database to Oracle database.
  • Used Git Version Control, JIRA for bug tracking, Trello for Dash boards and Confluence for project management.

Environment: Kinesis, Athena, Kafka, Boto3, EC2, S3, VPC, ELB, DynamoDB, Cloud Front, Cloud Watch, EC2, Security Groups, HSM, Red shift, CloudFormation, Spark, Spark SQL, AWS EMR, Hive, Apache, Sqoop, Python, PySpark, Airflow DAG, Linux, MySQL Oracle Enterprise DB, Jenkins, Oracle, Git, Oozie, MySQL.

Confidential, IL

Azure Data Engineer

Responsibilities:

  • Hands-on experience on Azure Cloud services, Management tools, Migration, Storage, Network &Content Delivery, Active Directory, Azure container service, VPN Gateway, Content Delivery Management, Azure Storage Services, Azure Database Services.
  • Developed Databricks pipelines to move the data from Azure blob storage/file share to Azure SQL Data warehouse and blob
  • Created CI/CD pipelines in Azure Devops using ARM templates to create VM’s, Resource Groups in Azure.
  • Written Scala and python script notebooks for Azure Databricks transformation task.
  • Worked on converting Hive/SQL queries into Spark transformations using Spark RDDs and Scala.
  • Worked on Spark Structured Streaming for developing Live Steaming Data Pipeline with Source as Kafka and Output as Insights into Cassandra DB where the data was fed in JSON/XML format and then Stored in Cassandra DB.
  • Configured VMs availability sets using Azure portal to provide resiliency for IaaS based solution and scale sets using Azure Resource Manager to manage network traffic.
  • Setup test environments in Windows Azure and worked on a plan to migrate applications to migrate to 100% Azure environments.
  • Created Azure Data Factory for copying data from Azure BLOB storage to SQL Server.
  • Embedded SSIS packages in Azure Data Factory using Azure-SSIS runtime
  • Created Pipeline in Azure Data Factory to connect to various azure resources like Blob Storage, Data Lake Storage, Azure SQL Database
  • Wrangled/Transformed data using Azure Data Factory Data Flows
  • Created various triggers in Azure Data Factory such as scheduling, tumbling, and storage event-based triggers
  • Used Azure Data Factory to execute Azure Synapse Spark Pool Notebook
  • Created Cache Memory on Windows Azure to improve the performance of data transfer between SQL Azure and WCF services.
  • Created a Virtual Network on Windows Azure to connect all the servers.
  • Worked on Azure Sql databases and azure functions to fulfill business requirements.
  • Creating Spark clusters and configuring high concurrency clusters using Azure Databricks to speed up the preparation of high-quality data.
  • Implement ETL process to move data from Cosmos to SQL Azure Database using SQLizer, SSIS, and SQL Azure Database
  • Set up deployment agents to deploy from Azure Devops to Azure.
  • Building and Installing Servers through Azure Resource Manager Templates or Azure Portal.

Environment: data bricks, PySpark, HDInsight, spark, sql, azure ADW, hive, Blob storage, Data Lake, APGateway Services, SSIS packages, RDBMS, SQL, Azure data lake, python scripts, ETL, Azure synapse Analytics

Confidential

Hadoop Developer

Responsibilities:

  • Experience installing and configuring Hadoop ecosystem components like Hadoop, Map Reduce, HDFS, Hive, Sqoop, and Flume.
  • Created Hive tables on HDFS to store the data processed by Apache Spark on the Cloudera Hadoop Cluster in Parquet format.
  • Developed a multi-node cluster in designing the Data Lake with the Cloudera Distribution.
  • Worked on extending Hive functionalities by writing UDFs to transform data in HDFS.
  • Deployed Databases which are stored using PostgreSQL and experienced in performing CRUD operations to operate on data schemas and browse functions in Python scripts.
  • Hands-on experience in Normalization (1NF, 2NF) Denormalization techniques for optimizing RDBMS performance.
  • Developed Sqoop scripts to load and transform structures and semi-structured data into HDFS.
  • Created Hive snapshot tables and Hive Avro tables from data partitions stored on S3 and HDFS.
  • Responsible for loading Data pipelines from web servers using Sqoop, Kafka and Spark Streaming API Worked on building large data clusters and real-time streaming with Spark in a team.
  • Experience in scheduling the jobs using Oozie workflows.
  • Expertise developing custom SQL queries for Tableau to generate Tableau reports and dashboards.
  • Involved in writing Sqoop script with incremental loading of Hive external tables.
  • Experience in creating partitions and bucketing concepts in Hive.
  • Extracted data from parquet files and stored them in PySpark data-frame for analysis
  • DBMS developments include building data migration scripts using Oracle SQL LOADER.
  • Leveraged Power BI ALM toolkit to compare Power BI Schema and metadata differences
  • Used tabular editor to create Calculation Group to decrease redundancy in measures and make them dynamic
  • Made use of Power BI Dataflow for creating centralized source with Power Query M language
  • Utilized Power BI Deployment Pipeline to move reports, dataflow, and datasets into different environments
  • Advanced knowledge in entity and relationship extraction from unstructured data
  • Made use of Power BI modeling layer to create relationships between different dimensions and facts
  • Wrote complex PL/SQL queries using joins, Stored Procedures, Functions, Triggers, cursors, and indexes in Oracle DBMS.
  • Architected the Start Schema using the modeling layer by using different concepts like bridge table and aggregated table
  • Enhanced Power BI refresh time by using various techniques like Incremental refresh and Manage Aggregations
  • Worked on Jenkins Pipelines to build Docker containers and exposure in deploying the same to Kubernetes engine.
  • Created SQL objects like Store Procedures, views, tables, and views to refine the SQL logic

Environment: Hadoop, Hive, Sqoop, kafka, Spark, Flume, SQL, PostgreSQL, Tableau, Oracle SQL,Oozie, DBMS, Cloud Formation, Pl/SQL, Jenkins.

Confidential

Data Analyst

Responsibilities:

  • Used Agile Methodology to implement project life cycles of reports design and development.
  • Work with Tableau to create dashboards, stories, and other visual displays, to provide clear and insightful. Create tools and analyze data that help with team efficiencies, scalability, and profitability.
  • Identify relevant trends, do follow-up analysis, prepare visualizations in form of charts, graphs, and tables.
  • Developing SQL scripts to validate the databases tables and reports data for backend database and data warehouse testing.
  • Process documents through SQL Server Management Studio (SSMS) and UI integrated tools.
  • Involved in Design, Development and Support phases of Software Development Life Cycle (SDLC).
  • Implemented Data Exploration to analyze patterns and to select features using Python.
  • Installed SQL Server DB and power BI, moved customer data given in CSV format into SQL Server DB.
  • Executed data validation, data profiling, data defects cleanup, data improvement initiatives, data quality monitoring reports.
  • Worked on data profiling and data validation to ensure the accuracy of the data between the warehouse and source systems.
  • Utilized excel functionality to gather, compile and analyze data from pivot tables and created graphs/charts.
  • Created reports in SSRS with different type of properties like chart controls, filters, Interactive Sorting, SQL parameters etc.
  • Worked with Data Cleaning, Wrangling, Data Analytics, Modelling, Integration, Data Science, Critical Thinking, Problem-Solving, Communication and Presentation Skills.

Environment: Python, Pyspark, SQL Hadoop, SSMS, SSRS, Teradata, Tableau, UNIX, Excel.

Confidential

Python developer

Responsibilities:

  • Professional experience as a Python Developer, proficient coder in multiple languages and environments including Python, REST API and SQL.
  • Proficiency in Agile methodologies such as Extreme Programming, Waterfall Model and Test-Driven Development.
  • Proficient with various python libraries like SciPy, NumPy, Matplotlib, Pandas to enhance the performance throughout the SDLC
  • Developed web-based application using Django framework with python concepts.
  • Generated Python Django forms to maintain the record of online users.
  • Used Django APIs to access the database.
  • Involved in Python OOD code for quality, logging, monitoring, and debugging code optimization.
  • Hands-on work developing in SAS, SQL, Python, and Java with Eclipse for extraction patterns from very large datasets and transform data into an informational advantage for decision support.
  • Developed web applications in Django Framework model view control (MVC) architecture.
  • Worked on JSON based REST Web services and Responsible for setting up Python REST API framework and spring framework using DJANGO.
  • Used Python and Django creating graphics, XML processing, data exchange and business logic implementation.
  • Worked on development of SQL and stored procedures on MYSQL.
  • Responsible for debugging the project monitored on JIRA (Agile).
  • Created database using MySQL, wrote several queries to extract data from database.
  • Implemented responsive user interface and standards throughout the development and maintenance of the website using the HTML, CSS, JavaScript, Bootstrap, and JQuery.
  • Model View Control architecture is implemented using Django Framework to develop web applications.
  • Used Python and Django creating graphics, XML processing, data exchange and business logic implementation.
  • Performed troubleshooting, fixed, and deployed many Python bug fixes of the two main applications that were a main source of data for both customers and internal customer service team.
  • Placed data into JSON files using Python to test Django websites. Used Python scripts to update the content in the database and manipulate files.
  • Held meetings with clients and worked all alone for the entire project with limited help from the client.
  • Used HTML, CSS, JQuery, JSON and JavaScript for front end applications.
  • Participated in the complete SDLC process and connected continuous integration system with GIT version control repository and continually built as the check-in's come from the developer.
  • Performed and assisted in design, development and testing of predictive analytics models that include large data collection, data organization, text segmentation, categorization, summarization and topic modeling.
  • Participating in the adoption of the approach for the implementation of the projects
  • Experience in writing Sub Queries, Stored Procedures, Triggers, Cursors, and Functions on PostgreSQL database.
  • Responsible for designing, developing, testing, deploying and maintaining the web application. Communicating with department to set goals and business objectives Involved in working with Python open stock API's.

Environment: Python, Django, Java Script, HTML, CSS, JSON, JQuery, XML, MYSQL, GitHub, SQL and Windows, Agile, Waterfall.

We'd love your feedback!