Sr. Cloud Data Engineer Resume
Chicago, IL
SUMMARY
- Migrated multitude of on - prem Data Warehouses and a Data Lake into Cloud Data Warehouse, Snowflake.
- As part of D&A architecture team, worked on holistic assessment of current capabilities and future requirements for our cloud transition to Snowflake.
- Separated workloads for business use cases using Snowflake virtual warehouses, which helped teams to forecast the expenses and prepare well for any future requirements.
- Worked on building a custom ingestion framework using PySpark and Snowflake JDBC functionalities to seamlessly migrate the Data from on-prem platforms to Snowflake.
- Implemented CDC and automated the Data load into higher environments using streams and tasks.
- Worked on Python module to automate reading the Excel workbooks from share point/Box folder and load into Snowflake
- Worked on Python module to integrate Snowflake with Organization’s centralized API, helped in calling APIs and update snowflake objects using the response
- Collect, Ingest, Transform and Deliver Data using GCP and Snowflake technologies to provide faster, deeper insights.
- Leveraged SQL, External functions, and few other advanced features to streamline pipeline development with no additional clusters, services, or copies to manage.
- Understanding and navigating Data needs to facilitate instant Data availability.
- Extensively worked on performance tuning of ETL pipelines and created best practices documents
- Expertise in all Informatica Intelligent Cloud Services (IICS) components, configured different task types for data integration or ingestion from external systems
- Migrated Data pipelines into GCP’s Dataflow/DataProc and streamlined pipeline development processes using templates.
- Provided high quality data and version control for SQL scripts using dbt. Used dbt to achieve referential integrity and data uniqueness.
- Engineered and Administered Informatica platforms for Cloud Services, Big Data Management, Master Data Management, Data Integration and Data Quality.
- Worked on cross-cloud replication and failover of snowflake databases between AWS and GCP.
- Developed Data Marts using Data warehouse and created Data Integration workflows using disparate technologies.
- Set up Continuous Integration/Continuous deployment pipelines using Jenkins and Docker which automated deploying changes into higher environments.
TECHNICAL SKILLS
Programming/Scripting Languages: SQL, Python, Bash
ETL/ELT: Informatica PowerCenter, Spark
Databases: Oracle, IBM DB2, SQL Server, MySQL, PostgreSQL, Amazon Athena, Hadoop/HBase, Hive, MS Access.
Cloud Technologies: Amazon Web Services, Google Cloud Platform.
Data warehouses/ Data lakes: Snowflake, Big Query, Google Cloud Storage, S3, IBM Netezza, Teradata
DevOps: Git, Bitbucket, Docker, Jenkins, Terraform
Schedulers/Orchestration: Airflow, CTRL-M, Autosys
BI Tools: Tableau, Power BI & Domo.
Other Applications: Microsoft Office, MS Visio, Box, Share point
Methodologies: Agile, Waterfall and Test-Driven Development
PROFESSIONAL EXPERIENCE
Confidential, Chicago, IL
Sr. Cloud Data Engineer
Responsibilities:
- Transition of Marketing Data from multitude of on-prem Data platforms to Cloud Data Warehouse, Snowflake.
- Built an opportunity database for migration support service (MSS) used by MSS Planning App.
- Worked with Systems Architects to create optimal virtual warehouses on Snowflake for all the Marketing applications.
- Automated administrative tasks which made application team to own and meter their performance and costs.
- Creating clusters on DataProc to answer business questions using PySpark.
- Created a process to align Content recommendations on website to related accounts using Snowflake and GCP.
- Extract internet-based networking Data from transactional Databases and load into Snowflake to provide valuable insights on Network Operations.
- Created Strategic Insights and Analytics Data on Snowflake and ingested into Domo using a custom connector.
- Collected data from data sources like salesforce, excel and normalize it in Amazon Redshift using Snaplogic ETL tool.
- Created complex SQL queries, Stored Procedures as a prerequisite data set to create reports.
- Designed Tableau Dashboards and extracts meaningful insights from the reports from the data collected in Redshift to Data Science Team.
- Managed Tableau Server and usability of the dashboards.
- Monitored failed pipelines in Snaplogic and resolves the issues on daily basis.
- Provided analytical and problem-solving abilities to clients on the data.
- Created Reports in Qlik sense and delivered them.
- Scheduled reports using Nprinting.
- Created data mappings in Informatica ETL tool to extract data from different source files, transform the data using Filter, Lookup, Aggregator, Expression, Transformations and then loaded into data warehouse.
- Created a process to provide near real-time Data using Snowflake, which helped to report accurate information.
- Built a near-real time pipeline, using SNOWPIPE for ingestion, Streams, and Tasks to orchestrate the data flow
- Provided high quality data and version control for SQL scripts using dbt. Used dbt to achieve referential integrity tests.
- Build data pipelines using Informati Confidential and load into Snowflake warehouse from disparate source systems
- Extensively worked on performance tuning of ETL pipelines and created best practice documents
- Developed shell scripts and scheduled to help identify new opportunities and also age out older opportunities for migration support
- Developed Snowflake procedures and SnowSQL scripts, to calculate scores for assessment and re-assessment
Environment: Snowflake, GCP, Big Query, Informatica/IICS, MDM, Linux, Oracle, Python, Spark, Airflow, dbt, Domo, Tableau, Power BI, Hashi corp Vault.
Confidential, Schaumburg, IL
Data Engineer
Responsibilities:
- Engineered and Administered Informatica platforms for Cloud Services, Big Data Management, Master Data Management, Data Integration and Data Quality.
- Transition of Zurich North America Warehouse from on-prem Netezza to Cloud Data Warehouse, Snowflake.
- Proactive cost management - created virtual Warehouses in a way that application team own the responsibility to manage costs better.
- Created a module using PySpark for one-time migration of the Data Warehouse and Data Lake.
- Creating Data pipeline to answer business questions using Spark, and Hadoop.
- Collaborate with client technical team to architect solutions for Predictive Modelling and Data Analytics on Data Lake using Spark and Hadoop.
- Implementation of Risk Exposure Data Store requirements on Data Lake using Azure Data Factory, Spark, and shell scripting.
- Developed & Automated creation of final recommendations summary table that will help to evaluate the readiness
- Effectively used RBAC which enabled granular role separation to manage applications.
- Pipelining to support Data science team which includes complex features.
- Real-time Semi-Structured JSON Data is made available for analytics using SNOWPIPE and shell scripting.
- Effectively using multiple tools and technologies to maintain CI/CD.
- Integrated multiple source systems into MDM to provide 360 view of customer and distributor master Data.
Environment: Informatica/IICS, Snowflake, Netezza, Salesforce, MDM, Azure, Python, Spark, Hadoop, Hive, Airflow, Linux.
Confidential, Collegeville, PA
Data Engineer
Responsibilities:
- Engineered and Administered Informatica platforms for Cloud Services, Data Integration, Data Virtualization and Data Quality.
- Configured HA/Grid setup to achieve High Availability and resiliency for all the production Informatica platforms.
- Configured applications on Informatica cloud (ICS), lift & shift of PowerCenter repositories into cloud
- Created Data quality and Data virtualization services (IDS) to accomplish Data Profiling, Data standardization and cleansing.
- Performance tuning for long running ETL processes to help with the Data Warehouse deliverables, Identifying root cause of the issues in platform/infrastructure.
- Created SOPs for all the informatica best practices, guidelines, issues, and resolutions
- Worked closely with Informatica Technical support on Informatica related technical & environmental issues.
- Configured Disaster recovery environment for all the business-critical applications.
- Created UNIX shell scripts for automating regular maintenance activities using Informatica Command line utilities.
Environment: Informatica/ICS, Linux, Oracle, DB2 9.5, Teradata, Microsoft SQL Server
Confidential
Data Integration Developer
Responsibilities:
- Business requirements analysis, process flow diagrams of present and future state.
- Designed architecture for web analytics Data warehouse and data marts.
- Implemented ETL solution for new Data marts.
- Created mappings/workflows to extract Data from SQL Server, Sybase and Flat File sources and load into various Business Entities and Transaction Data Sets for Automatic Position Reporting.
- Source Data analysis and characterization of Data.
- Designed logical Data model and assist in physical modelling.
- Development of required oracle packages/procedure/functions.
- Creating technical specification documentations for the mappings developed.
Environment: Informatica, Oracle 11g, Unix, Perl, Business Objects.
