We provide IT Staff Augmentation Services!

Sr Azure Data Engineer Resume

0/5 (Submit Your Rating)

CA

SUMMARY

  • A data professional with 8 years of progressive experience in ETL, Pipeline Building, Data Analytics, Data Modelling, Reporting, Visualization and. Excellent capability in collaboration, quick learning and adaptation, effective interpersonal and communication skills, proven leadership skill and proficient decision - making capabilities.
  • Used various Hadoop distributions (Cloudera, Hortonworks, Amazon EMR, Microsoft Azure HDInsight) to fully implement and leverage new Hadoop features.
  • Hands-on experience in developing and deploying enterprise-based applications using major Hadoop ecosystem components like MapReduce, YARN, Hive, HBase, Flume, Sqoop, Spark MLlib, Spark SQL, and Kafka.
  • Adopt at configuring and installing Hadoop/Spark Ecosystem Components.
  • Worked with Spark to improve efficiency of existing algorithms using Spark Context, Spark SQL, Spark MLlib, Data Frame, Pair RDD's and Spark YARN.
  • Experience in application of various data sources like Oracle SE2, SQL Server, Flat files, and unstructured files into a data warehouse.
  • Able to use Sqoop to migrate data between RDBMS, NoSQL databases and HDFS.
  • Experience in Extraction, Transformation and Loading (ETL) data from various sources into Data Warehouses, as well as data processing like collecting, aggregating, and moving data from various sources
  • Worked on ETL Migration services by developing and deploying AWS Lambda functions for generating a serverless data pipeline which can be written to Glue Catalog and can be queried from Athena.
  • Extract, Transform and load data from sources systems to Azure Data Storage services using a combination of Azure data factory, T-SQL, Spark SQL. Data ingestion to one or more Azure services (Azure Data Lake, Azure storage, Azure SQL) and processing data in Azure Data bricks.
  • Experience in implementing Azure data solutions, provisioning storage account, Azure Data Factory, SQL server, SQL Databases, SQL Data warehouse, Azure Data Bricks and Azure Cosmos DB.
  • Hands-on experience with Hadoop architecture and various components such as Hadoop File System HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Hadoop MapReduce programming.
  • Designed and Developed ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
  • Ample knowledge of data architecture including data ingestion pipeline design, Hadoop/Spark architecture, data modeling, data mining, machine learning and advanced data processing.
  • Experience working with NoSQL databases like Cassandra and HBase and developed real-time read/write access to very large datasets via HBase.
  • Developed Spark Applications that can handle data from various RDBMS (MySQL, Oracle Database) and Streaming sources.
  • Fluent programming experience with Python, SQL, T - SQL, PySpark for querying, data extraction/transformations and developing queries for a wide range of applications.
  • Strong analytical and problem-solving skills and the ability to follow through with projects from inception to completion.
  • Ability to work effectively in cross-functional team environments, excellent communication, and interpersonal skills.

PROFESSIONAL EXPERIENCE

Sr Azure Data Engineer

Confidential, CA

Responsibilities:

  • Working on data management disciplines including data integration, modeling, and other areas directly relevant to business intelligence/business analytics development.
  • Working on Microsoft Azure toolsets including Azure Data Factory Pipelines, Azure Data bricks, Azure Data Lake Storage.
  • Experience on Migrating SQL database to Azure Data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and controlling and granting database access and Migrating On premise databases to Azure Data Lake store using Azure Data factory.
  • Create and maintain optimal data pipeline architecture in cloud Microsoft Azure using Data Factory and Azure Databricks.
  • Building the pipelines to copy the data from source to destination in Azure Data Factory.
  • Applied spark streaming for real time data transforming.
  • Implemented a generic ETL framework with high availability for bringing related data for Hadoop & Cassandra from various sources using spark.
  • Experience in building data pipelines using Azure Data factory, Azure Databricks and loading data to Azure Data Lake, Azure SQL Database, Azure SQL Data warehouse and controlling and granting database access.
  • Working on query languages such as SQL, code languages such as Python or C# and scripting languages such as PowerShell, M-Query (Power Query), or Windows batch commands.
  • Creating data visualizations using Power BI.
  • Enhance & maintain the data warehouse in snowflake to support all reporting BI needs
  • Working on internal and external aggregated pipelines to make them available to end users.
  • Working on Azure (IAAS), Should be involved in the planning, design, and deployment of Cloud solutions.
  • Installed, configured, administered, monitored Azure, IAAS and PAAS, Azure AD.
  • Experience working in deployment scripts in PowerShell.
  • Responsible for developing, support and maintain the ETL processes using Informatica PowerCenter
  • Develop and maintains the Resource groups and instances.
  • Extensive Experience of designing, developing, and deploying various kinds of reports using SSRS using relational and multidimensional data.
  • Experience with ad-hoc reporting, Parameterized, Custom Reporting using SSRS for daily reports.
  • Azure VMs, Networking (VNet’s, Load Balancers, App Gateway, Traffic Manager, etc.)
  • Provided high availability for IaaS VMs and PaaS role instances for access from other services in the VNet with Azure Internal Load Balancer.
  • Implemented and tested python-based web applications interacting with MySQL.
  • Experience with Microsoft Azure, Azure Resource Management templates, Virtual Networks, Storage, Virtual Machines, and Azure Active Directory.

Environment: Microsoft Azure (Data Lake, Data Lake Analytics, SQL Database, Powershell, Data Bricks and SQL Data warehouse), MySQL, SQL, Python, PySpark, PowerBI, ETL, Informatica

Data Engineer

Confidential, Columbus, OH

Responsibilities:

  • Partnered with ETL developers to ensure that data is well cleaned, and the data warehouse is up to date.
  • Selected and generated data into CSV files and stored them into AWS S3 by using AWS EC2 and then structured and stored in AWS Redshift.
  • Experience with POC that involves scripting using PySpark in Azure Data bricks.
  • Data Extraction, aggregations, and consolidation of Adobe data within AWS Glue using PySpark.
  • Worked with AWS cloud platform and its features which include EC2, IAM, EBS CloudWatch and AWS S3
  • Deployed application using AWS EC2 standard deployment techniques and worked on AWS infrastructure and automation. Worked on CI/CD environment on deploying application on Docker containers.
  • Used AWS S3 Buckets to store the file and injected the files into Snowflake tables using Snow Pipe and run deltas using Data pipelines.
  • Experience in moving data between GCP and Azure using Azure Data Factory.
  • Used PySpark and Pandas to calculate the moving average and RSI score of the stocks and generated them into data warehouse.
  • Generated report on predictive analytics using Python and Tableau including visualizing model performance and prediction results.
  • Developed python script to transfer data from on-prem to S3 and script to hit REST API's and extract data to S3.
  • Worked on Ingesting data by going through cleansing and transformations and leveraging AWS Lambda, AWS Glue and Step Functions.
  • Extensively used ETL for supporting data extraction, transformations and loading processing, in a complex EDW using Talend/Data stage.
  • Created YAMLfiles for each data source and including glue table stack creation.
  • Worked on a python script to extract data from Netezza databases and transfer it to AWS S3.
  • Developed Lambda functions and assigned IAM roles to run python scripts along with various triggers (SQS, Event Bridge, SNS).
  • Created a Lambda Deployment function and configured it to receive events from S3 buckets.
  • Writing UNIX shell scripts to automate the jobs and scheduling cron jobs for job automation using commands with Crontab.
  • Developed various Mappings with the collection of all Sources, Targets, and Transformations using Informatica Designer.
  • Developed Python scripts to update content in the database and manipulate files.
  • Developed Mappings using Transformations like Expression, Filter, Joiner and Lookups for better data messaging and to migrate clean and consistent data.
  • Exported the analyzed data into relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Converting data load pipeline algorithms written in python and SQL to spark and PySpark.
  • Mentor and support other members of the team (both on-shore and off-shore) to assist in completing tasks and meet objectives.
  • Utilized Agile and Scrum methodology for team and project management.
  • Used Git for version control with colleagues. CI/CD pipeline has been used to deploy the code to Production.

Environment: AWS (Lambda, Glue, S3, Step functions, Redshift, EC2), SQL, Python, PySpark, CI/CD, Git, ETL, Snowflake, API’s, Tableau

Data Engineer

Confidential, Fort Worth, TX

Responsibilities:

  • Gathered business requirements and prepared technical design documents, target to source mapping document, mapping specification document.
  • Extensively worked on Informatica PowerCenter.
  • Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregations to build the data model and persists the data in HDFS.
  • In depth Knowledge of AWS cloud service like Compute, Network, Storage, and Identity & access management.
  • Designed AWS architecture, Cloud migration, Dynamo DB and event processing using Lambda function
  • Managed storage in AWS using Elastic Block Storage, S3, created Volumes and configured Snapshots
  • Gathered requirements from Business and documented for project development.
  • Coordinated design reviews, ETL code reviews with teammates.
  • Developed mappings using Informatica to load data from sources such as Relational tables, Sequential files into the target system
  • Developed automation system using PowerShell scripts and JSON templates to remediate the Azure services.
  • Parsed complex files through Informatica Data Transformations and loaded it to Database.
  • Optimized query performance by oracle hints, forcing indexes, working with constraint-based loading and few other approaches.
  • Extensively worked on UNIX Shell Scripting for splitting group of files to various small files and file transfer automation.
  • Worked with Autosys scheduler for scheduling different processes.
  • Assisted in UAT Testing and provided necessary reports to the business users. Performed unit testing to validate data flow from end to end.
  • Developed Python and SQL scripts to extract data from various databases and Spark code using Spark-SQL for faster testing and data processing.
  • Involved in development, testing and postproduction for the entire migration project.

Environment: AWS, SQL, Python, Spark, ETL, Powershell, Unix, Informatica, HDFS, Dynamo DB, S3, Automation, Autosys, UAT testing

ETL Developer

Confidential

Responsibilities:

  • Implemented reporting Data Warehouse with online transaction system data.
  • Worked with PL/SQL procedures and used them in Stored Procedure Transformations.
  • Extensively worked on oracle and SQL server. Wrote complex SQL queries to query ERP system for data analysis purpose
  • Deployed mircoservices2, including provisioning AZURE environment.
  • Developing pipelines in Azure Data Factory using SQL Azure.
  • Created SSIS Packages for moving data from Excel Files and SQL Server to data warehouse.
  • Migrating current data center environment to Azure Cloud using tools like Azure Site Recovery (ASR).
  • Build Data Sync job on Windows Azure to synchronize data from SQL 2012 databases to SQL Azure.
  • Expert in implementing the in-memory computing capabilities like Apache Spark written in Pyspark
  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics.
  • Worked in converting Hive/SQL queries into Spark transformations using Spark RDDs.
  • Experience in building power bi reports on Azure Analysis services for better performance.
  • Microsoft Azure IaaS Services, Active Directory Domain Services (ADDS), DNS, SSL, PowerShell scripting.
  • Tuned ETL jobs in the new environment after fully understanding the existing code.
  • Maintained Talend admin console and provided quick assistance on production jobs.
  • Involve in designing Business Objects universes and creating reports.
  • Built ad hoc reports using stand-alone tables.
  • Experience in Azure Marketplace where to search, deploy and purchase wide range of applications and services.
  • Wrote Custom SQL for some complex reports.
  • Performed analysis after requirements gathering and walked team through major impacts.
  • Provided and debugged crucial reports for finance teams during month end period.
  • Addressed issue reported by Business Users in standard reports by identifying the root cause. Get the reporting issues resolved by identifying whether it is report related issue or source related issue.

Environment: Microsoft Azure, SQL (Spark SQL, T-SQL, PL/SQL), Python, PySpark, ETL, Data Analytics, ADF

Data Analyst

Confidential

Responsibilities:

  • Involved in understanding the legacy applications & data relationships.
  • Attended user design sessions, studied user requirements, completed detail design analysis, and wrote design specs.
  • Utilized AWS services with focus on big data analytics, enterprise data warehouse and business intelligence solutions to ensure optimal architecture, scalability, and flexibility.
  • Interacted with key users and assisted them with various data issues, understood data needs and assisted them with Data analysis.
  • Implemented and tested python-based web applications interacting with MySQL.
  • Developed the Pyspark code for AWS Glue jobs.
  • Worked on scalable distributed data system using Hadoop ecosystem in AWS, MapR distribution.
  • Build cluster on AWS environment using S3, EC2, Redshift.
  • Prepared and maintained documentation for on-going projects.
  • Worked with Informatica PowerCenter for data processing and loading files.

Environment: AWS, SQL, Python, Glue, S3, EC2, Redshift, PySpark, MYSQL

We'd love your feedback!