Sr. Data Modeler/data Engineer Resume
Jersey City, NY
SUMMARY
- Above 10+ years of IT experience in Data Modeling, Data Engineering and Data Analysis with high proficiency in using big data technologies using cloud infrastructure,
- Expert in Agile and Software Development Life Cycle (SDLC) and expertise in detailed design documentation.
- Excellent experience with Azure Data Lake, Azure Data Factory, Azure Synapse Analytics, Azure Data Bricks and AWS EC2&S3.
- Expertise in designing Star schema, Snowflake schema for Data Warehouse, ODS architecture by using tools like Erwin data modeler, Power Designer, and ER/Studio.
- Experience on data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining and advanced data processing.
- Expert in writing SQL queries and optimizing the queries in Oracle, DB2 and Teradata.
- Experience in handling Big Data using Hadoop eco system components like Sqoop, Pig and Hive.
- Hands on experience on Python programming PySpark implementations in AWS EMR, building data pipelines infrastructure to support deployments for Data Analysis and cleansing
- Expertise in scheduling JAD (Joint Application Development) with End Users, stake Holders, Subject Matter Experts, Developers and Testers.
- Enterprise in Data Management Data Governance, Data Modeling, Warehousing, Data Integration, Meta - data, Reference Data and MDM.
- Experience in SQL, PL/SQL package, function, stored procedure, triggers, and materialized view, to implement business logics of oracle database.
- Experience in using Business Intelligence tools (SSIS, SSAS, and SSRS) in Tableau.
- Good understanding and hands on experience in setting up and maintaining NoSQL Databases like HBase.
- Excellent experience on TeradataSQLqueries, TeradataIndexes, Utilities such as Mload, Tpump, Fastload and FastExport.
- Strong experience in Normalization(1NF, 2NF, 3NF and BCNF) and De-normalization techniques for effective and optimum performance
- Experience in developing TSQL scripts and stored procedures to perform various tasks and multiple DDL, DML, and DCL activities to carry out business requirements
- Expertise in writing Stored Procedures, Functions, Nested Functions, building Packages and developing Public and Private Sub-Programs using PL/SQL and providing Documentation.
- Efficient in Extraction, Transformation and Loading (ETL) data from spread sheets, database tables using Confidential data transformation service (DTS)
- Strong RDBMS concepts and well experience in creating database Tables, Views, Sequences, triggers, joins taking the Performance and Reusability into consideration.
- Excellent knowledge in Data Analysis, Data Validation, Data Cleansing, Data Verification and identifying data mismatch
- Heavy use of Access queries, V-Lookup, formulas, Pivot Tables, etc.
- Solid experience in building ODS, EDW, Staging, data mart, semantic layer and maintaining metadata repository.
- Excellent experience in writing SQLqueries to validate data movement between different layers indatawarehouse environment.
- Strong experience in using MS Excel and MS Access to dump the data and analyze based on business needs.
TECHNICAL SKILLS
Data Modeling Tools: Erwin 9.8, Sybase Power Designer, ER/Studio V17
Big Data tools: Hadoop3.0, HDFS, Hive2.3, Kafka1.1, HBase1.2, Sqoop1.4.
Database Tools: IBM DB2, Oracle 12c/11g, Teradata15/14, and MS Access.
Azure Cloud Platform: Azure Data Lake Gen2, Azure Data FactoryV2, Azure Data Bricks, Azure Synapse Analytics, Azure CosmosDB.
AWS Cloud Platform: Amazon Redshift, AWS Ec2&S3, AWS Lambda, AWS Glue, AWS Athena.
Project Execution Methodologies: Ralph Kimball and Bill Inmon data warehousing methodology, JAD, Agile.
Reporting tools: SQL Server Reporting Services (SSRS), Tableau, Crystal Reports, Business Objects
ETL Tools: SSIS, Informaticav10.
Programming Languages: SQL, T-SQL, and PL/SQL.
Operating Systems: Confidential Windows 10/8/7, UNIX, and Linux
Tools: & Software: TOAD 6.2, MS-Office suite (Word, Excel, MS Project and Outlook)
PROFESSIONAL EXPERIENCE
Confidential
Sr. Data Modeler/Data Engineer
Responsibilities:
- Worked as Sr. Data Modeler/Data Engineer to review business requirement and compose source to target data mapping documents.
- As a lead used to suggest, maintain, modify Claims data model for improvements and optimization.
- Extensively used agile methodology as the Organization Standard to implement the data Models.
- Maintained source and target mappings, transformation logic and processes to reflect the changing business environment over time
- Created a Data Model that can store Producer related data coming from three different source systems onto a single Producer Data Catalog Model.
- Loaded data into Hive Tables from Hadoop Distributed File System (HDFS) to provide SQL access on Hadoop data
- Lead a team of 7 data modelers and address data model repository maintenance, exclusive locking, and ETL Development issues.
- Created Data factory pipelines that can bulk copy multiple tables at once from relational database to Azure data lake gen2
- Created various types of data visualizations using Python
- Deployed Tableau Connect with Azure SQL Synapse Analytics to set up the data source.
- Used Azure Cosmos DB for storing catalog data and for event sourcing in order processing pipelines.
- Created a Data Warehouse Model with de-normalized attributes at all places.
- Built a naming standards file from scratch to help in consistency.
- Conducted JAD sessions with management, vendors, users and other stakeholders for open and pending issues to develop specifications.
- Involved in Data Architecture, Data profiling, Data analysis, data mapping and Data architecture artifacts design.
- Guided the full lifecycle of a Hadoop solution, including requirements analysis, platform selection, technical architecture design, application design and development, testing, and deployment
- Built a IBM DB2 data model for the Cosmos system
- Created DDL scripts using Erwin and source to target mappings to bring the data from source to the warehouse.
- Migrated on-premise Oracle ETL process to Azure Synapse Analytics.
- Authored Python (PySpark) Scripts for custom UDF’s for Row/ Column manipulations, merges, aggregations, stacking, data labeling and for all Cleaning and conforming tasks.
- Developed Databricks Python notebooks to Join, filter, pre-aggregate, and process the files stored in Azure data lake storage Gen2.
- Identified the facts dimensions grain of fact, aggregate tables for Dimensional Models
- Implemented Source-to-Target mapping with the necessary transformation logic for the ETL team to build
- Involved in design and development of multiple Power BI Dashboards and reports.
- Done data migration from an RDBMS to a NoSQL database, and gives the whole picture for data deployed in various data systems.
- Imported and exported data using Sqoop from HDFS to RDBMS and vice-versa
- Worked with Azure BLOB and Data lake storage and loading data into Azure SQL Synapse analytics (DW).
- Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System to pre-process the data.
- Improved performance and optimization of the existing algorithms, explored different components like Spark Context, Spark-SQL, accumulators.
- Worked on Data governance, data quality, data lineage establishment processes.
- Ensured ETL succeeded and loaded data successfully in SnowflakeDB.
- Studied the source data model to come up with the necessary joins and filter conditions to get the right data in the target database
- Worked with Azure Blob Storage, Azure Data Lake Gen2, Azure Data FactoryV2, Azure SQL, Azure SQL Datawarehouse, Azure Synapse Analytics and AzureDatabricks.
- Gained Comprehensive knowledge and experience in process improvement, normalization/de-normalization, data extraction, data cleansing, data manipulation
- Worked in helping to create a Swagger API file for the application to consume with JSO
- Installed, configured and maintained Data Pipelines
- Implemented ETL process to move data from Cosmos to SQL Azure Database using SQLizer and SQL Azure Database.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Designed and developed user defined functions, stored procedures, triggers for Azure Cosmos DB
- Handled importing data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
- Handled performance requirements for databases in OLTP and OLAP models.
- Worked with Azure Data Warehouse, Azure Storage Accounts, Azure Data Factories, Azure DataLake Gen2, U-SQL and Blob storage.
- Written SQL Queries, Dynamic-queries, sub-queries and complex joins for generating
- Designed and implemented PL/SQL store procedures, functions and packages for data manipulation and validation.
- Developed the Logical and physical data model and designed the data flow from source systems to Oracle tables and then to the Target system.
- Leveraged Python development environment for data analysis and report building.
- Tested the ETL process for both before data validation and after datavalidation process.
- Extensively used Kafka and integrating with the Spark Streaming
- Worked in azure Data Lake storage gen2 to store excel files, parquet file.
- Involved in design, development, and validation & testing of Data warehouses using ETL and Data Modeling.
- Worked with DBA to create the physicalmodel and tables.
Environment: Erwin9.8, IBM DB2, Agile, Azure, SnowflakeDB, ETL, Hadoop3.0, Hive2.3, SQL, PL/SQL, Oracle12c, Power BI, Python, Spark, Azure CosmosDB, PySpark, OLAP, OLTP, HDFS.
Confidential - Jersey City, NY
Data Modeler/Data Engineer
Responsibilities:
- Implemented logical and physical relational database and maintained Database Objects in the data model using Erwin
- Implemented cloud data lakes like Azure Data Lake Gen2.
- Worked on migration of data from On-prem SQL server to Cloud databases (Azure Synapse Analytics (DW) & Azure SQL DB).
- Created and Configured Azure Cosmos DB.
- Worked in Agile environment and participated in daily Stand-ups/Scrum Meetings.
- Developed a data pipeline using Kafka to store data into HDFS.
- Ingested huge volume and variety of data from disparate source systems into Azure Data Lake Gen2 using Azure Data Factory V2.
- Monitored the scheduled Batch jobs for the execution of Loading Process in MDM.
- Worked on maintaining and managing Tableau and POWER BI driven reports and dashboards.
- Designed and implemented database solutions in Azure SQL Data Lakes, Azure Data Lake, Azure Data Factory, Azure Synapse Analytics, Azure SQL
- Created dimensional model for the reporting system by identifying required dimensions and facts using Erwin.
- Data sources are extracted, transformed, and loaded to generate CSV data files with Python programming and SQL queries.
- Developed Map Reduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW
- Architected and documented Azure SQL Data Warehouse (Synapse Analytics).
- Worked with relational database systems (RDBMS) such as Oracle and database systems like HBase.
- Wrote Python normalizations scripts to find duplicate data in different environments.
- Used Reverse Engineering to connect to existing database and create graphical representation.
- Used SQL on a wide scale for analysis, performance tuning and testing
- Developed Data Mapping, Data Governance, Transformation and Cleansing rules for the MasterData Management Architecture involving OLTP, ODS and OLAP.
- Designed and Developed Oracle, PL/SQL Procedures for Data Import/Export and Data Conversions.
- Deployed the DDL in the Development environment and required database
- Exported the analyzed data to the relational databases using Sqoop for visualization.
- Effectively worked in Azure Synapse Analytics, Azure Data Lake and Azure Data Bricks.
- Coordinated with DBA on data base build and table normalizations and de-normalizations
- Configured Input & Output bindings of Azure Function with Azure Cosmos DB collection to read and write data from the container whenever the function executes.
- Defined Big Data strategy, including designing multi-phased implementation roadmaps.
- Analyzed Azure Data Factory and Azure Data Bricks to build new ETL process in Azure.
- Used Star Schema and Snowflake Schema methodologies in building and designing the Logical Data Model into Dimensional Models
- Used Azure Data Factory as an orchestration tool for integrating data from upstream to downstream systems.
- Created several types of data visualizations using Python and Tableau.
- Implemented Spark Scripts using Scala, Spark SQL to access hive tables into spark for faster processing of data.
- Involved in writing T-SQL working on SSIS, SSAS,DataCleansing,DataScrubbing andDataMigration.
Environment: Erwin9.8, Agile, SQL, Oracle12c, PL/SQL, Big Data3.0, HDFS, OLAP, OLTP, ODS, Sqoop1.4, Azure Data Lake, ETL, SSIS, SSAS, Hive2.3, DBA.
Confidential - Columbus, GA
Sr. Data Modeler
Responsibilities:
- Designed both 3NF data models for ODS, OLTP systems and dimensional data models using star and snowflake Schemas.
- Designed the business requirement collection approach based on the project scope and SDLC methodology.
- Involved with Business Analysts team in requirements gathering and in preparing functional specifications and changing them into technical specifications.
- Created data models for AWS Redshift and Hive from dimensional data models.
- Transferred the data using Informatica tool from AWS S3 to AWS Redshift.
- Conducted statistical analysis on Healthcare data using python and various tools.
- Created logical and physical data models using Erwin and reviewed these models with businessteam and data architecture team.
- Worked with Data governance and Data quality to design various models and processes.
- Used SQL Server Integrations Services (SSIS) for extraction, transformation, and loading data into target system from multiple sources
- Developed Advance PL/SQL packages, procedures, triggers, functions, Indexes and Collections to implement business logic using SQL Navigator.
- Worked with AWS to implement the client-side encryption as Dynamo DB does not support at rest encryption at this time.
- Experience used ERWIN forward engineer to generate schema in ORACLE, SQL/SERVER environments.
- Used Teradatautilitiesfastload, multiload, tpump to load data
- Worked with AWS Lambda using python to automate resource creation, perform compliance checks and cost optimization.
- Worked in importing and cleansing of data from various sources like Oracle, flat files, with high volume data
- Implemented Installation and configuration of multi-node cluster on Cloud using Amazon WebServices (AWS) on EC2.
- Implemented Data Archiving strategies to handle the problems with large volumes of data by moving inactive data to another storage location that can be accessed easily.
- Created reports using SQL Reporting Services (SSRS) for customized and ad-hoc Queries
- Used AWS Glue for the data transformation, validate and data cleansing.
- Created Hive queries that helped analysts spot emerging trends by comparing fresh data with EDW reference tables and historical metrics.
- Created external tables with partitions using AWS Athena and Redshift
- Maintained metadata (data definitions of table structures) and version controlling for the datamodel.
- Tested the messages published by ETLtool and data loaded into various databases
- Developed the performance tuning of the database by using EXPLAIN PLAN, TKPROF utilities and also debugging the SQL code.
Environment: Erwin, SQL, Teradata, Amazon Redshift, Oracle, PL/SQL, SSRS, Hive, AWS, ETL, SSIS, AWS Athena.
Confidential - Bellevue, WA
Data Modeler/Data Analyst
Responsibilities:
- Performed the Data Mapping, Data design (Data Modeling) to integrate the data across the multiple databases in to EDW.
- Extracted the data from AWS RedShift into HDFS using Sqoop.
- Performed Data Modeling, Database Design, and Data Analysis with the extensive use of ER/Studio.
- Extensively performed Data analysis using Python Pandas.
- Conducted data modeling JAD sessions and communicated data related standards.
- Created Conceptual, Logical and Physical data models.
- Pulled data daily from OLTP and database to OLAP or staging database using the SSIS packages.
- Developed complex T-SQL code such as Stored Procedures, functions, triggers, Indexes, and views for the business application.
- Created Rich dashboards using Tableau Dashboard and prepared user stories to create compelling dashboards to deliver actionable insights.
- Processed the data using HQL (like SQL) on top of Map-reduce.
- Worked in importing and cleansing of data from various sources like DB2, Oracle and flat files.
- Worked with data compliance teams, Data governance team to maintain data models, Metadata, Data Dictionaries define source fields and its definitions.
- Organized User Acceptance Testing (UAT) conducted presentations and provided support for Business users to get familiarized with Loan products application.
- Troubleshooted test scripts, SQL queries, ETL jobs, and Enterprise data warehouse/datamart/data store models.
- Implemented a proof of concept deploying this product in Amazon Web Services (AWS).
- Created summary tables using de-normalization technique to improve complex Join operations
- Performed Data modeling for existing Databases using Toad Data Modeler.
- Worked on designing the whole data warehouse architecture from scratch, from ODS to datamarts
- Created partitions and indexes for the tables in the data mart.
- Worked on creating role playing dimensions, fact less Fact, snowflake and star schemas.
- Validated the data of reports by writing SQL queries in PL/SQL Developer against ODS.
- Developed SQLscripts for loading data from staging area to target tables.
- Designed the dimensional Data Model of the data warehouse.
- Configured report server and authorized permissions to different users in SQL Server Reporting Services (SSRS)
Environment: ER/Studio, DB2, AWS, Tableau, Sqoop, UAT, SQL, PL/SQL, ODS, T-SQL, OLAP, OLTP.
Confidential, Highland park, NJ
Data Analyst/Data Modeler
Responsibilities:
- Worked on data analysis, data profiling, data modeling, data mapping.
- Developed ER and Dimensional Models using Power Designer advanced features.
- Worked on SQL Server Integration Services (SSIS) to integrate and analyze data from multiple heterogeneous information sources.
- Design of Redshift Data model, Redshift Performance improvements/analysis
- Involved in Migrating the data model from one database to Teradata database.
- Stored the data with dictionary data type in Python, which the key-value pair could indicate a random point as key and any of its neighbors as value.
- Developed dimensional model for Data Warehouse/OLAP applications by identifying required facts and dimensions.
- Worked on AWS, implementing solutions using services like (EC2, S3, RDS and Redshift)
- Interacted with business users to understand the business requirements.
- Generated tableau dashboards for sales with forecast and reference lines
- Developed SQLscripts for creating tables, Sequences, Triggers, views and materialized views.
- Participated in meetings, reviews, and user group discussions as well as communicating with stakeholders and business groups.
- Extensively completed data quality management using information steward and did extensive data profiling.
- Developed Staging jobs where in using data from different sources like flat files, excel files, Oracle database
- Created Informatica mappings using various Transformations like Joiner, Aggregate, Expression, Filter and Update Strategy.
- Developed triggers, stored procedures, functions and packages using cursors and ref cursor concepts associated with the project using Pl/SQL
- Built reports and report models using SSRS to enable end user report builder usage.
- Designed and developed the data dictionary and Meta data of the models and maintain them.
- Translated logical data models into physical database models, generated DDLs for DBAs
- Enforced referential integrity in the OLTP data model for consistent relationship between tables and efficient database design.
- Used T-SQL queries to pull the data from disparate systems and Data warehouse in different environments.
- Performed unit testing, system integrated testing for the aggregate tables.
Environment: Power Designer, SQL, Oracle, AWS, Teradata, PL/SQL, T-SQL, OLTP, Python, Informatica, DBA.
Confidential
Data Analyst
Responsibilities:
- Performed data analysis on the target tables to make sure the data as per the business expectations.
- Created customized report using OLAP Tools such as Crystal Report for business use.
- Performed Data Validation and Data Reconciliation between disparate source and target systems for various projects.
- Worked with data investigation, discovery and mapping tools to scan every single data record from many sources.
- Involved in extensive Data validation by writing several complexSQL queries.
- Developed regression test scripts for the application.
- Worked closely with the SSISDevelopers to explain the complex Data Transformation using Logic.
- Managed timely flow of business intelligence information to users.
- Migrated critical reports using PL/SQL&UNIX packages.
- Created and scheduled the job sequences by checking job dependencies.
- Involved in metrics gathering, analysis and reporting to concerned team and tested the testing programs.
- Created or modified the T-SQL queries as per the business requirements.
- Generated various reports using SQL Server Report Services (SSRS) for business analysts and the management team.
- Wrote complex SQLqueries for validating the data against different kinds of reports generated by Business Objects.
- Created Column Store indexes on dimension and facttables in the OLTPdatabase to enhance read operation.
- Used advanced Confidential Excel to create pivot tables.
- Developed re-usable components in Informatica, and UNIX.
- Created ad-hoc reports to users in Tableau by connecting various data sources
Environment: SQL, PL/SQL, UNIX, OLAP, T-SQL, SSIS, SSRS, Excel, OLTP, Informatica, Tableau.
