Senior Data Integration Specialist Resume
Eden Prairie, MN
PROFESSIONAL SUMMARY:
- Over 9+ years of working experience in various stages of software development, out of which 2.5 years senior level experience in full life cycle (SDLC) Data Warehousing, Data Integration, Business Intelligence as Senior Data Integration Specialist, Data Warehouse Architect using IBM Infosphere, Websphere DataStage 11.3, 8.7,8.5,7.5, SQL, DB2 UDB 8/9, Oracle10g/9i/8i, PLSQL, Shell Scripts, Unix, Linux, Windows.
- 6+ months experience in Bigdata Hadoop components and Talend ETL tool.
- Proficiency in defining ETL design and integrating data from disparate source systems and disparate data formats including Requirements gathering, Source System Analysis, Data Quality, Slowly Changing Dimensions mappings, ETL - QA, Performance Tuning, Maintenance and Production Support using IBM DataStage 11.3,8.7,8.5, 7.5, Bigdata Hadoop Hbase, Hive and Talend ETL Tool and Custom-built scripts using PLSQL and Unix Korn Shell
- Strong experience in Data Modeling and deployment of multi-dimensional schema (Star and Snowflake Schemas), Conceptual/Logical/Physical design.
- Strong programming and debugging skills in SQL, PL/SQL and UNIX shell scripting. Experience in developing Stored Procedures, Functions, Packages, Database triggers, Materialized views and extensive experience in Query optimization, Performance tuning
- Experience with Teradata, DB2, Oracle 10g and DB tools TOAD, Teradata SQL Assistant, SQL Dbx
- Team leadership, Onsite/Offshore activities and exposed to driving the translation and construction of a client's complex business problems into innovative technology solutions. Ability to learn very quickly and apply new skills to existing problems.
TECHNICAL SKILLS:
ETL: Talend, DataStage 11.3, 8.7, 8.5 and 7.5
Data Modeling: ERWIN 4.x, Visio
Database tools: Hbase, Hive, TOAD, SQL Dbx, Quest Central for DB2, Quest Central for Toad, Teradata SQL Assistant, service now, IBM Control Center, Command Editor
RDBMS: DB2, Oracle 8i/9i/10g/11g, Teradata
Languages: Python, HTML5, CSS, JavaScript, UNIX Shell scripting, C, PL/SQL
O.S: UNIX (HP-UX, AIX), Linux, Windows7/Windows XP/2000/ NT/9x
Tools/Utilities: Tivoli Workload Scheduler and Broker (TWSD & TWSB), Autosys 4.5, MS Visio
PROFESSIONAL EXPERIENCE:
Confidential, Eden Prairie, MN
Senior Data Integration SpecialistEnvironment: IBM DataStage 11.3, Talend 6.3, Teradata 15.0, Big Data, Hadoop, Linux, Service Now, IBM DB2, Tivoli workload Scheduler/broker, CLM Migration tool
Responsibilities:
- End-to-end implementation, maintenance, optimizations and enhancement of the application.
- This involves various phases in the project, starting from project estimation, Requirement Gathering and Analysis, Designing, Coding, User Acceptance Testing to deployment in production, including post implementation support. Documenting the Proof of Concepts and delivered to Clients. Writing of Technical Specification of the project.
- Design activities for Batch Control Architecture, Mapping Relationships, Keys and Indexes
- ETL Development for Control Architecture, Common Modules, Sequence Controls and major critical interfaces.
- Developed Talend jobs to import data from Datalake and load it into tables.
- Executing Talend jobs and debugging the errors.
- Designed Parallel jobs using various stages like Join, Remove Duplicates, FTP stage, Filter.
- Extensively writing SQL, PL/SQL scripts performing DDL operations.
- Working under agile methodology, as part of this crated user stories, tasks and participating in daily scrum calls.
- Created job definitions Job streams and migration list to migrate jobs to test and Production environments.
- Used Quality Stage for various data cleansing stages to get complete visibility of the actual condition of data, to reformat data from multiple systems to ensure that the data has the correct specified content and format, and to ensure that the best available data survives and is correctly prepared for the target.
- Worked with multiple database structures and performed extraction of legacy data and load into Teradata target database
- Profiling and minimizing overall costs and resources for critical data integration projects by scanning the samples of data and determining their quality and structure.
- Scheduling jobs using the Tivoli work load scheduler.
- Create a Data Lake on the HDFS/Hadoop platform and perform data analytics and data ingestion from traditional and non-traditional big data sources using native Hadoop tools - Habse, Hive etc.
- Used development/debug stages to test the environment by creating samples of data from given high volume data or by creating mock data. Ensuring timely deliveries of work items to the Client
- Worked on the code fixes and on the tickets raised due to the job failures.
- Supporting unit, integration, and end user testing by resolving identified defects.
- Involved in Implementing ETL standards and Best practices within our portfolio.
- Used the Datastage Designer to develop processes for extracting, cleansing, transforming, integrating, and loading data into data warehouse database. Reusing the logic from Datastage jobs in real time
- Do schedule cleanup and zombie process cleanup.
- Take ownership of production failures and resolve them ASAP
- Work on Data anomaly issue tickets
- Work on problem tickets
Confidential, Eden Prairie, MN
Data Warehouse ArchitectEnvironment: IBM DataStage 8.5, Teradata 14.10, Linux, Db2, IBM DB2, Tivoli workload Scheduler
Responsibilities:
- Involved with Application and Database Support Teams to maintain system and database which supports current data scenarios and easily adapts to changing business needs.
- Used Parallel Extender extensively by using Processing, Development/Debug, File and Real Time Stages.
- Involved in all phases including Requirement Analysis, Design, Coding, Testing, Support and Documentation
- Extensively used Datastage Designer to develop processes for extracting, transforming, integrating and loading data from various sources into the Data Warehouse database.
- Extensively used Qualitystage stages like Standardize, Match frequency, duplicate & unduplicated match and Survive stages.
- Widely Used different types of DataStage stages like modify, sequential file, Copy, Aggregator, Surrogate key, Transformer, dataset, look up, join, Remove Duplicates, sorter, Column generators, CDC, and Funnel.
- Extracted data from Heterogeneous source systems like Oracle, Teradata, SQL Server and flat files.
- Developed processes on both Teradata and Oracle using shell scripting and RDBMS utilities such as Multi Load, Fast Load, Fast Export, BTEQ (Teradata) and SQL*Plus, SQL*Loader (Oracle).
- Developed Multi load and Fast Load scripts to load data from flat files into Teradata Staging area.
- Used database stage like DB2/UDB Enterprise, DB2 bulk load, ODBC Connector
- Extensively Oracle Enterprise and Oracle Connector stages
- Worked on Integration testing of Application using Converted data to check for any discrepancies.
- Involved in Generating an SOA service of DataStage, QalityStage job and deploy/Manage the services using IBM Console to receive service requests.
- Worked with Data Modelers, Technical Architects, Customer/end user, Business analysts, and Data analysts to design Technical Specification documents / Mapping documents.
- Extensively tested all the developed jobs based on the requirements of the project and deployed the code into production environment.
- Used Quality Stage plug-in stage to call Quality Stage jobs into the DataStage Designer.
- Extensively Implemented slowly changing dimensions (SCD) TYPES.
- Created and used Shared Containers to reuse the logic in another Job.
- Developed staging and Data Mart DS jobs using Data Stage Designer on parallel environment.
- Extensively used SQL in Datastage jobs for processing data.
- Improved development in job sequences for the designed jobs using exec command Job Activity, Triggers and E-mail Notification Activities.
- Experience in writing different complex types of SQL’s for validating the data between different stages with in the data warehouse.
- Very good team player and excellent communication skills.
Confidential, Eden Prairie, MN
ETL Developer
Environment: IBM DataStage 8.7, 7.5, DB2, TOAD for Db2, Tivoli workload Scheduler
Responsibilities:
- Developed DataStage server jobs to extract, transform and load data into data Warehouse from various sources like relational databases (DB2), Oracle 9i, flat files etc.
- Worked with Business Analysts to analyze the business requirements and functional specifications.
- Used Parallelism concepts for distributing load among different processors by implementing Pipeline and partitioning of data. Involved in Designing Parallel and server Jobs.
- Interpreted logical and physical data models for Business users to determine data definitions and establish referential integrity of the system.
- Involved in creating the projects. Improved application performance by tuning SQL statements and fixing proper indexes, Designed data models.
- Extensively used Parallel Stages like Join, Merge, Lookup, Filter, Remove Duplicates, Funnel, Row Generator, Modify, Peek etc. for development and de-bugging purposes
- Extensively worked with Data Stage Job Sequences to Control and Execute Data Stage Jobs and Job Sequences using various Activities and Triggers.
- Used Data Stage Director and the runtime engine to schedule running the server jobs, monitoring scheduling and validating its components.
- Scheduled the parallel jobs using DataStage Director, which is controlled by DataStage engine and also for monitoring and performance statistics of each stage.
- Created data quality standardization jobs using Web Sphere Quality Stage and also by writing PL/SQL queries to identify and analyze data anomalies, patterns, inconsistencies etc.
- Responsible for metadata management, new job categories and creating new data elements, creating shared containers for reusability.
- Worked on performance tuning to address very critical and challenging issues.
- Implemented the Surrogate Key by using Key Management functionality for newly inserted rows in Data Warehouse.
- Responsible for daily verification that all scripts, downloads, and file copies were executed as planned, troubleshooting any steps that failed, and providing both immediate and long-term problem resolution.
- Developed UNIX scripts to automate the Data Load processes to the target Data warehouse.
- Created Error Tables containing data with discrepancies to analyze and re-process the data.
- Used DataStage Director and its run-time engine to schedule running the solution, testing and debugging its components, and monitoring the resulting executable versions (on an ad hoc or scheduled basis).
- Worked with DataStage Manager for importing metadata from repository, new job Categories and creating new data elements.
- Generations of Surrogate IDs for the dimensions in the fact table for indexed, faster data access.
- Implemented shared containers to use in multiple jobs, which have same business logic.
- Interaction with the business users to better understand the requirements and document their expectations, handling the current process, modifying and created the jobs to the updated requirements, handle the load process to data mart and eventually data warehouse.
- Involved in the design, development and testing of the PL/SQL stored procedures, packages and triggers for the ETL processes.
- Defined and implemented approaches with Metadata Definitions, Import and Export of Datastage jobs using Datastage tools functionality. Test critical jobs in DEV and migrate into TEST env for UAT.
- Prepared test cases and complete testing with in time.
- Do a parallel run testing in DataStage 7.5 and DataStage 8.7 versions
Confidential, Eden Prairie, MN
ETL DeveloperEnvironment: IBM DataStage 7.5, Oracle 10g, SQL, PL/SQL, TOAD 11, Tivoli workload Scheduler
Responsibilities:
- Senior Integration/DataStage Analyst and Data Architect
- Actively participate in understanding business requirements, analysis and designing Data Migration/Integration Process from Business Analyst
- Translate business requirements into Data Migration architectural design
- Create Technical Design document, Source to Target mapping documents based on the client business requirement
- Define foundational areas such as policies and procedures, standards, testing framework and an operational model applicable to any required to DataStage implementation project
- Design and implement Meta data repository and reusable components
- Create ETL best practices and Developers guide
- Use DataStage to extract, transform and load data from all Applications (SAP BW, SAP BPC, Enterprise Data Warehouse) to Staging and Target
- Create reusable containers, scripts and DS jobs to be used in multiple transformations
- Develop the strategy Flat file to Relational table, table to table load mappings
- Handle the performance tuning of DataStage mappings
- Responsible to test the DataStage (mappings)
- Responsible for code migration, Code review, test plans, test scenarios, test cases as part of Unit/Integrations testing
- Automate the Production jobs using Tivoli workload scheduler, Autosys
Confidential
ETL DeveloperEnvironment: IBM DataStage 7.5, Oracle 10g, SQL, PL/SQL, TOAD 11, Tivoli workload Scheduler
Responsibilities:
- Responsible for designing, developing, implementing all aspects of ETL Phase
- Defined star schema and staging database environments
- Developed high performance ETL jobs to load large volumes of data
- Designed and developed ETL jobs to load data from DB2 UDB relational tables
- Developed various Complex DataStage jobs including slowly changing dimensions
- Designed and implemented slowly changing dimension types
- Developed various Complex DataStage Mappings implementing SCD type 2 for Shipment, Purchase Order, Receipt, Location, Item Dimensions etc., and also implemented incremental load strategy for Fact table load (Purchase Item, EDI Match & Receipt Item)
- Designed bulk loading utility for large volume of source data
- Detected bottlenecks, trouble shooting and tuning of mappings, queries, sessions for better performance and efficiency. Handled error handling routine using automated error checking mechanisms to reduce failures & monitor error logs
- Created ETL batch jobs and sequencers for data warehouse data loading automation process. Designed a strategy to integrate workflows with Control-M scheduler
- Created UNIX Shell scripts to automate routine tasks
- Deliver project documentations such as Source to target mapping document, ETL design specs, ETL Test strategy, Production Operational documents
Confidential
ETL DeveloperEnvironment: IBM DataStage 7.5, Db2, IBM AIX, Shell script, SAS, Autosys 4.5, IBM Db2, MS Office tools.
Responsibilities:
- Involved in ‘Public Sector’ and EDW projects for the analysis and understand the Source systems (FACETS, COSMOS, SRCL and UNET).
- Involved in meetings with the Business Analysts to collect the requirements, analysis and implementation of it and prepared Specification documents for EDW process.
- Created logical and physical Dimensional data models using Erwin
- Developed Data marts for users as per their requirements.
- Coordinate with team members, distributed the specifications and designs and monitor work assigned.
- As part of EIW (Enterprise Information Warehouse) standards coding involved analysis, cleansing, transformation and deliverance across the ETL environment.
- Designed and developed various Datastage Parallel jobs to extract, transform and loading the data into EDW and DM tables.
- Used shared containers for multiple jobs, which have the same business logic.
- Migrated the jobs from Datastage 7.5 to 8.1 and used Multi-Client manager to work in different versions.
- Developed custom build ops to implement business rules and complex logic based on requirement.
- Created Shared Containers to increase Object Code Reusability and to increase thru put of the system
- Developed test plans, test cases and performed unit testing of jobs to ensure that it meets the requirements and quality.
- Developed Perl and Korn shell and tested the Unix Shell Scripts used in loading the data into the EIW and Data Marts
- Scheduled all the jobs using Autosys to automate all daily/monthly jobs.
- Created various reports according to the user requirements.
- Created MS Visio Process flows to understand the processes and for maintenance purpose.
- Prepared implementation plans and check lists for installs/code migration.
- Using CVS all migration will be done by checking in the code within UNIX and migrating across different environments from Development to SIT (System Integration testing environment), UAT (User Acceptance testing environment) and finally to Production.
- Used PAC2000 V 7.1 tool for creating Work Orders for production installs and Problem Tickets for Production issues.
- Coordination User Acceptance Testing and providing end user trainings.
- Supported Production On call on rotation basis, for newly installing code and for all Production issues.
Source Systems: Purchase Order Management System, EDI856, EDI214, 3D, ART, Info Retriever and Central Receipt Database.
