Data Engineer Resume
Dallas, TX
SUMMARY
- Around 3 years of IT experience in the areas of Data Science, Data mining, Data Analysis, Software design, development, and testing in Bigdata and Cloud Ecosystem. Passionate in learning new tools and technologies in Cloud Computing
- Excellent understanding of Hadoop Ecosystem including HDFS, Map Reduce, Hive, Kafka, Spark, YARN, HBase, Oozie, and Sqoop based Bigdata Platforms.
- Very good understanding of Cloud (both Azure and GCP) Ecosystem including ADF, Azure Data lake, Azure Data Bricks, and Google Big query.
- Develop End to end Orchestration and alerts using Azure Data Factory and Azure Event hub
- Expertise in Hive Query Language (HQL), Hive security, and debugging Hive issues
- Experience with infrastructure setup using Terraforms and ARM template on Azure cloud platform (both IaaS and PaaS)
- Very good understanding of the Azure Resources and Services like VMs, Containers, Blob Storage, Azure Hyper - scale services, Key Vault, App Services, and Traffic Manager.
- Created Azure security policies to restrict users in accessing resources in Azure
- Created Azure Clusters for High Availability application performance.
- Good knowledge of key DevOps concepts - Continuous Integration and Delivery.
- Designed Azure virtual machines (VMs) and VM architecture for IaaS and PaaS; understand the availability
- Responsible for performing extensive data validation using Hive Dynamic partitioning and Bucketing
- Worked on a different set of tables like internal and external tables on Hive
- Complete knowledge of Data Warehouse Methodologies Star and Snowflake Schemas of Dimensional Modeling and Slowly Changing Dimensions (SCD's).
- Involved in the entire life cycle of Data warehouse design, development, and Implementation.
- Proficient in gathering requirements and authoring Business Requirement Documents (BRD) into System Requirement Specifications (SRS) and identifying interface and business process specifications.
- Proficient in developing Entity-Relationship diagrams, Star/Snowflake Schema Designs, and Expert in modeling Transactional Databases and Data Warehouse.
- Excellent understanding of ETL process, Data modeling (Dimensional & Relational) on concepts like Star-schema, Snowflake schema using fact and dimension tables and relational databases and client/server applications.
- Extensively involved in Reverse-Engineering and Business rules extraction expert in project planning, strategic planning, systems analysis and troubleshooting, quality control, forecasting, scheduling, and planning, and tracking of results.
- Experience in integration of various data sources with Multiple Relational Databases like Oracle, SQL Server, MS Access and Worked on integrating data from flat files.
- Good experience in Agile Engineering practices, Scrum methodologies, and Test-Driven Development and Waterfall methodologies.
- Good knowledge of various scripting languages like Linux/Unix, shell scripting, and Python.
- Extensively worked on Spark Streaming and Apache Kafka to fetch live stream data.
- Experience in converting Hive/SQL queries into RDD transformations using Apache Spark, Scala, and Python.
- Experience in data processing like collecting, aggregating, moving from various sources like Kafka.
TECHNICAL SKILLS
Programming Languages: SQL, Java, MapReduce, Python, C Programming, HTML, Cloud Functions
Database Management: Big Data Hadoop, Hive, HBase, MySQL, GCP Big query, Oracle
Software: Microsoft Office (Word, Excel, PowerPoint), Jupyter Notebook, Tableau, GCP, GitHub, IBM Watson, putty and Cloudera Manager, Hue, Azure Data bricks, ADLS, and Azure Synapse
Operating Systems: iOS (all versions), Mac OS, Windows, Linux
PROFESSIONAL EXPERIENCE
Confidential
Data Engineer
RESPONSIBILITIES:
- Constructed a Data ingestion pipelines to process semi-structured and relational data from various data sources to Google Cloud Platform
- Maintained data Pipelines up to 99.8% while ingesting streaming and transactional data across various data sources using python and Prefect.
- Ingested data from disparate data sources using a combination of SQL, and Google Analytics API using Python to create data views to be used in BI tools like Tableau.
- Created multiple code snippets of DBT for Data Transformation.
- Migration from GitHub to DBT, while automating the ETL process across billions of rows of data.
- Communicated with project managers and analysts about data pipelines that drove efficiency.
- Build basic ETL that ingested historical and transactional data from a Web API with about 12,000 daily active users.
- Analyzed data of developer in the SecAPI where the issues are about a million records in a day and ingested this data in Big Query through the Google Storage Buckets.
- Worked with clients to understand needs of business and translate those business needs into actionable reports in Tableau
- Project Documentation and Production Deployment
- Participated in daily SCRUM calls to update progress, raise blockers, and identify next task to work.
- Worked with event based google cloud functions which trigger the ingestion pipelines when there is new data in the storage buckets.
- Used Kanban framework to implement agile and DevOps software development which requires real time communication and full transparency of work.
Environment: Python, SQL, Cloudera Manager, Google Cloud Functions, Prefect, DBT, Google Big Query, UNIX, Shell Scripting, GCP, and Tableau
Confidential, Dallas, Tx
Cloud Developer
RESPONSIBILITIES:
- Created Data ingestion pipelines to ingest data from HDFS and RDBMS to Azure Data lake services
- Build ingestion Pipelines on using ADF to Azure Data lake from external RDBS and flat files ingestion
- Created Infrastructure in azure using Terraform.
- Created multiple code snippets of Terraform to create Infrastructure in Azure.
- Created Azure policies to restrict users to create and access resources in Azure.
- Hands-on Experience on DevOps pipelines.
- Hands-on Experience in Azure services.
- Developed End to end Orchestration and alerts using Azure Data Factory and Azure Event hub
- Build Business Models using Python with Jupyter Notebook on Azure Data Bricks
- Analyzed data and create pictorial graphs using Python libraries and Tableau for Business reporting.
- Real-time streaming the data using Spark on Azure Data Bricks
- Involved in source system analysis, data analysis, data modeling to ETL (Extract, Transform, and Load).
- SIT and UAT support for the Developed Applications
- Project Documentation and Production Deployment
- Participated in daily SCRUM calls to update progress, raise blockers, and identify next task to work.
- Collaborate with cross-functional teams to define, design, and implement new code Modules
- Used Agile methodology to emphasize face-to-face communication and passing the iteration through a full software development cycle.
Environment: Azure Databricks, ADLS, Azure Synapse, Apache Hadoop 2.2.0, Cloudera Manager, MapReduce, Hive, Hue, HDFS, Sqoop, Kafka, UNIX, Shell Scripting, GCP, Bigquery, and Tableau
Confidential, Dallas, Tx
Cloud Developer- Intern
RESPONSIBILITIES:
- Worked with Google studios to perform visualization of data stored on GCP
- Build ingestion Pipelines on using ADF to Azure Data lake from external RDBS and flat files ingestion
- Created Infrastructure in azure using Terraform.
- Created multiple code snippets of Terraform to create Infrastructure in Azure.
- Created Azure policies to restrict users to create and access resources in Azure.
- Developed End to end Orchestration and alerts using Azure Data Factory and Azure Event hub
- Build Business Models using Python with Jupyter Notebook on Azure Data Bricks
- Analyzed data and create pictorial graphs using Python libraries and Tableau for Business reporting.
- Real-time streaming the data using Spark on Azure Data Bricks
- Involved in source system analysis, data analysis, data modeling to ETL (Extract, Transform, and Load).
- Used Agile methodology to emphasize face-to-face communication and passing the iteration through a full software development cycle.
Environment: Azure Databricks, ADLS, Azure Synapse, Apache Hadoop 2.2.0, Cloudera Manager, MapReduce, Hive, Hue, HDFS, Sqoop, Kafka, UNIX, Shell Scripting, GCP, Bigquery, and Tableau
Confidential, New York
Cloud Developer
RESPONSIBILITIES:
- Created Data ingestion pipelines to ingest data from HDFS and RDBMS to GCP Bucket
- Performed joins, group by, and other aggregate operations as per Business rules on GCP Bigquery
- Build a custom process to convert call center data (speech to text translation) to tables using Google Speech to Text Converter
- Build ingestion Pipelines on using GCP functions to GCP Buckets from external RDBS and flat files ingestion
- Created Infrastructure in GCP using Terraform.
- Worked with Google studios to perform visualization of data stored on GCP
- Created multiple code snippets of Terraform to create Infrastructure on GCP.
- Created Azure policies to restrict users to create and access resources in GCP.
- Hands-on Experience on DevOps pipelines.
- Hands-on Experience in Google services including Google Functions.
Environment: Python, UNIX, Shell Scripting, GCP, GCP Bigquery, Google functions, Google Studios and Tableau
Confidential
Bigdata/Azure Developer - Intern
RESPONSIBILITIES:
- Involved in analyzing the ETLs on Netezza/Hadoop and supported for the solution to migrate the Data/processes from IBM Netezza and Hadoop to Azure.
- Assisted in building the ETL source to Target specification documents by understanding the business requirements
- Developed Spark scripts and UDFs for analyzing the data loaded in Azure from various data sources to ADLS using Azure Data bricks and created external tables on ADLS for Aggregate layer set up on Synapse
- Implements various Hive optimization techniques like bucketing, map-side joins, bucketed map-joins, converting to parquet file storage format, etc
- Used Sqoop to import data into HDFS from Oracle, MySQL, Netezza, and Access databases and vice-versa.
- Created HBase tables to store variable data formats of data coming from different applications.
- Experience in managing and reviewing Hadoop log files.
- Created a custom input format in map-reduce to read the excel files that are imported into HDFS and convert them into CSV format.
- Developed and executed shell scripts to automate the jobs
- Worked with several file formats such as Text files, Sequence files, RC files, Avro files, ORC files, Parquet files, Custom INPUT FORMAT, and OUTPUT FORMAT in Spark and Hive.
- Channelized MapReduce outputs based on requirement using Practitioners.
- Created Oozie workflows for Spark, MapReduce, Hive and Sqoop, SSH, and Shell actions.
- Create Oozie coordinators to schedule the workflows based on time dependency and data dependency.
- Coordinated and worked closely with legal, clients, third-party vendors, architects, DBA’s, operations and business units to build and deploy.
- Prepared Test Strategy and Test Plans for Unit, SIT, UAT, and Performance testing.
ENVIRONMENT: Apache Hadoop, Cloudera Manager, Azure Databricks, ADLS, Synapse, MapReduce, Hive, Hue, Hbase, HDFS, Sqoop, Oozie, Netezza, Putty, Winscp.
