We provide IT Staff Augmentation Services!

Sr. Data Modeler (hadoop) Resume

0/5 (Submit Your Rating)

Minnetonka, MN

SUMMARY

  • Sr. Data Modeler with 6+ years of experience in data analysis and modelling for Online Transactional processing (OLTP),Data warehousingand Online Analytical Processing (OLAP)systems.
  • Experience in various phases of Software Development life cycle (Analysis, Requirements gathering, Designing)
  • Experience with JAD sessions for requirements gathering and writing functional specifications, queries.
  • Good experience in Relational and Dimensional Data modeling for creating Logical and Physical Design of Database and ER Diagrams using multiple data modeling tools like ERWIN, ER Studio and Power - Designer.
  • Deep knowledge in Normalization / De normalization techniques for optimum performance in relational and dimensional database environments
  • Well versed in Data Warehousing concepts like Star Schema and Snowflake Schema, Slowly changing dimensions, foreign key concepts, referential integrity
  • Hands on experience with modeling using ERWIN in both forward and reverse engineering processes and created DDL scripts for implementing Data Modeling changes.
  • Have hands on experience in writing Map Reduce (with Java and Python) jobs on Hadoop Ecosystem and implementing UDF’s in HIVE and PIG.
  • Experience in using Teradata Utilities such as T pump, Fast Load and M load. Created BTEQ scripts.
  • Extracting Mega Datafrom Amazon Redshift and Elastic Searchengine using SQL Queries to create reports.
  • Adept at writing Data Mapping Documents, Data Transformation Rules and maintaining Data Dictionary.
  • Experience in Managing scalable Hadoop clusters including Cluster designing, provisioning, custom configuration
  • Knowledge in architecting Hadoop solutions including hardware recommendations, network topology design, storage configurations, benchmarking, performance tuning, administration and support.
  • Coordinated with other developers in tuning long running SQL queries to enhance system performance.
  • Created naming convention files and worked with DBAs to create the physical model and tables.
  • Proficient in writing complex SQL Queries, Sub Queries, views, stored procedures, Normalization, Database Design, Functions, Triggers
  • Excellent knowledge in Data Analysis, Data Validation, Data Cleansing, Data Verification and involved in designing and Preparing Test Scenarios, Test Plans, Test Cases and TestData.
  • Expertise in Hadoop eco system components HDFS, Map Reduce, Yarn, HBase, Pig, Sqoop, Flume and Hive for scalability, distributed computing and high-performance computing.
  • A strong team member with excellent analytical, interpersonal and communication skills.

TECHNICAL SKILLS

Modeling Tools: Erwin r9.7, Sybase Power Designer 16.6, Oracle Designer, ER/Studio.

Database Tools: Microsoft SQL Server 2017, Teradata 15, Oracle 12c, MS Access 2016, Poster SQL, Netezza.

OLAP Tools: Tableau, SAP BO, SSAS, Business Objects, and Crystal Reports.

ETL Tools: SSIS, Informatica Power Center, Web Intelligence, SSRS.

Operating System: Windows 10/8, Dos, Unix

Big Data: Hadoop, HDFS, Map Reduce, Hive, Pig, Impala, Sqoop, Flume, Oozie, Kafka, Spark - Scala, PySpark

Reporting Tools: Business Objects, Crystal Reports SAP SE, SAP Business Intelligence, Micro Strategy.

Web technologies: HTML 5, DHTML, XML.

Tools: & Software's: Toad, MS Office 2016, BTEQ, Elasticsearch, Teradata SQL Assistant.

Other tools: SQL*Plus, SQL*Loader, MS Project, MS Visio 2016 and MS Office 2016

PROFESSIONAL EXPERIENCE

Confidential, Minnetonka, MN

Sr. Data Modeler (Hadoop)

Responsibilities:

  • Worked as a Data Modeler to generate Data Models using Erwin and subsequent deployment to Enterprise Data Warehouse.
  • Participated in JAD sessions for defining business requirements and finalizing the required data fields and formats
  • Developed logical data models and physical database design and generated database schemas using Erwin persists high volume data predictive analytics
  • Installed and configured Hadoop, Map Reduce, HDFS (Hadoop Distributed File System), developed multiple MapReduce jobs in java for data cleaning.
  • Coordinated withDBA to create best-fit PhysicalDataModel from the Logical Data Model using Erwin
  • Generated DDL (DataDefinition Language) scripts using Erwin and supported the DBA in Physical Implementation ofdataModels.
  • ImplementedElasticSearchand Log Stash stack to collect and analyze the logs produced by the spark cluster.
  • Reverse Engineered the existing Stored Procedures and wrote Mapping Documents for them.
  • Used Erwin editors to modify, update existing primary, foreign, alternate keys and identifying, non-identifying, recursive and subtype relationships.
  • Automated workflows using shell scripts pull data from various databases into Hadoop
  • Deployed Hadoop Cluster in Fully Distributed and Pseudo-distributed modes.
  • Performed tuning and optimization of complex SQL queries using Teradata Explain.
  • Used Teradata Data Mover to copy data and objects such as tables and statistics from one system to another.
  • Performed data analysis and data profiling using SQL queries on various sources systems including Oracle, SQL Server and DB2.
  • Involved with the coders in evaluation of CPT codes to ensure that the diagnosis meets medical necessity for the specific CPT code.
  • Modified Oracle PL/SQL codes like stored Procedures, Functions, Triggers etc. based on technical and functional specification documents. Used Sub-queries, Merge statements and Joins extensively in stored procedures.
  • Defined ETL framework which includes load pattern for staging and ODS layer using ETL tool, file archival process, data purging process, batch execution process.
  • Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
  • SearchAPI will haveElasticSearchcapability.
  • Assisted in designing test plans, test scenarios and test cases for integration, regression and user acceptance testing.

Environment: MS Visio, SQL Server 2014, Hadoop,Java, MapReduce, HDFS, Hive, Spark- Scala, Sqoop, Kafka, Java (jdk1.6), PL/SQL, Agile,ETL, Oracle 11g, ElasticSearch, DB2, Teradata Erwin.

Confidential, Rancho-Cordova, CA

Sr. Data Modeler (Hadoop)

Responsibilities:

  • Translated business requirements into detailed, production-level technical specifications, new features, and enhancements to existing technical business functionality.
  • Analysis of functional and non-functional categorized data elements for Data Migration,
  • Data profiling and mapping from source to target data environment.
  • Designed database models for the operational data store, Data Warehouse, and federated databases to support client enterprise Information Management Strategy.
  • Having solid involvement in utilizing Teradata utilities like TPT, FASTLOAD, MULTILOAD, and BTEQ scripts.
  • Excellent Programming and optimization experience in Teradata SQL.
  • Shared responsibility for administration of Hadoop, Hive and Pig.
  • Designed the DataMart’s using Ralf Kimball’s Dimension DataMart modelling methodologies using Erwin.
  • Extensively used Star Schema methodology for the reporting system by identifying required dimensions and facts and cleansed unwanted tables/columns.
  • Worked on Big Data Hadoop cluster implementation and data integration in developing large-scale system software.
  • Installed and configured MapReduce, HIVE and the HDFS; implemented CDH3 Hadoop cluster on Centos. Assisted with performance tuning and monitoring.
  • Designed logical data model, physical data model, data mapping and ETL design.
  • Created logical/physical models and conceptual models using ERwin and VISIO.
  • Normalized the tables up to 3NF and established referential integrity of the system.
  • Created partitions and indexes for the tables in thedatamart.
  • Involved in designing and implementing the security for the databases.
  • Maintained metadata (datadefinitions of table structures) and version controlling for thedatamodel.
  • Performed data analysis, Data Migration and data profiling using complex SQL on various sources systems including Oracle and Teradata.
  • Developed component for connecting ElasticSearchDB.
  • Created Source to Target mappings and Transformations. Mapped data between Source and Targets.
  • Wrote SQL scripts to test the mappings and Developed Traceability Matrix of Business Requirements mapped to Test Scripts to ensure any Change Control in requirements leads to test case update.
  • Developed data transformation and cleansing rules for migration using ETL tools

Environment: Erwin, Teradata, Oracle,SQL ServerErwin, Hadoop, Map Reduce, Hive, HDFS, PIG, Sqoop, Oozie, Data lake, Agile, ETL, Agile, ElasticSearch, MS Visio,Teradata Crystal Reports, Windows XP.

Confidential, Woonsocket, RI

Data Analyst/Data Modeler

Responsibilities:

  • Processed large Excel source data feeds for Global Function Allocations and loaded the CSV files into Oracle Database with SQL Loader utility.
  • Performed code Inspection and moved the code into Production Release.
  • Performed Data filtering, Dissemination activities, trouble shooting of database activities, diagnosed the bugs and logged them in version control tool.
  • Used ElasticSearch for creating reports like Blocked events Report and Handler Execution reports for providingdataon processed and unprocessed events in the system.
  • Created Logical Data Models and Physical Data Models using Erwin Data Modeler.
  • Created the conceptual, Logical and Physical Model for the data warehouse with emphasis on insurance (life and health), mutual funds and annuity using Erwin data modeling tool.
  • Analyzed the processes in medical coding and transition.
  • Performed source data analysis and captured metadata, reviewed results with business. Corrected data anomalies as per business recommendation.
  • Designed star schema with dimensional modeling, created fact tables and dimensional tables.
  • Involved in data analysis, data discrepancy reduction in the source and target schemas.
  • Designed and developed the star schema Data model, Fact Tables to load the Data into Data Warehouse.
  • Implemented one-many, many-many Entity relationships in the data modeling of Data warehouse.
  • Analyzed the tables once the imports were done.
  • Perform database tuning and optimize complex SQL queries usingTeradataExplain, stats and indexes.
  • Created physical and logical data models using Erwin.
  • Created source to target mappings for multiple source from SQL server to oracle.
  • Performed the batch processing of data, designed the SQL scripts, control files, batch file for data loading.
  • Performed the Data Accuracy, Data Analysis, Data Quality checks before and after loading the data.
  • Designed and developed the Database objects (Tables, Materialized Views, Stored procedures, Indexes), SQL statements for executing the Allocation Methodology and creating the OP table, CSV, Text files for business.
  • Performed the physical database design, normalized the tables, worked with Denormalized tables to load the data into fact tables of Data warehouse.
  • Worked with Repository Manager, Designer, Workflow Manager and Monitor to import and load Source Definitions using Source Analyzer and Target Definitions using Warehouse Designer.
  • Complex mappings using corresponding Source, Targets and Transformations like update strategy, lookup, stored procedure, SQL, sequence generator, joiner, aggregate, Java and expression transformations in extracting data in compliance with the business logic.
  • Created SSIS Packages to migrate slowly changing dimensions.

Environment: SQL Server, DTS,SSIS, SQL Server Data Tools,Data lake, Agile, Visual Studio, Visual Source Safe,Erwin, Team Foundation System.

Confidential

Data Analyst/Data Modeler

Responsibilities:

  • Facilitated JAD sessions for project scoping, requirements gathering & identification of business subjectareas.
  • Identified and documented detailed business rules and use cases based on requirements analysis.
  • Developed Logical Model from the conceptual model.
  • Worked with DBA's to create a best-fit Physical Data Model from the logical data model.
  • Developed dimensional model for Data Warehouse/OLAP applications by identifying required facts and dimensions.
  • Fix Teradata utility failures and ability to handle other Mainframe, Teradata related errors by making necessary code overrides.
  • Extensively used the Erwin design tool & Erwin model manager to create and maintain the versions of the Inland Marine data model.
  • Performance Tuning (Database Tuning, SQL Tuning, Application/ETL Tuning)
  • Design and direct the implementation of security requirements for the data warehouse and direct the information access and delivery effort for the data warehouse.
  • Direct the data warehouse Meta data capture and access effort and also define Meta data standards for the data warehouse.
  • Load data into Teradata from legacy systems and flat files using MultiLoad scripts and FastLoad scripts.
  • Write complex SQL queries to pull the required information for business use from database using TeradataSQL Assistant.
  • Designed STAR schema for the detailed data marts and plan data marts consisting of confirmed dimensions.
  • Identified and documented data sources and transformation rules required to populate and maintain data Warehouse content.
  • Involved in extracting, cleansing, transforming, integrating and loading data into different Data Marts using Data Stage Designer.
  • Developed SQL Queries to fetch complex data from different tables in remote databases using joins, database links and bulk collects.
  • Used Erwin's reverse engineering to connect to existing database and ODS to graphically represent the Entity Relationships.
  • Involved in Performance tuning by leveraging oracle explain utility and SQL tuning.

Environment: Oracle SQL Server DB2, Microsoft Transaction Server, Erwin, Teradata Crystal Reports, Windows XP.

We'd love your feedback!