Sr. Etl/informatica Developer Resume
Chicago, IL
SUMMARY
- Over 9+ years of IT experience on Client - Server architecture, Business Intelligence applications in various phases like analysis, design, deployment, coding and testing using Talend, Informatica, and SSIS on OLAP and OLTP environments.
- Extensively used ETL methodology for performing DataProfiling, DataMigration, Extraction, Transformation and loading using Talend/SSIS/Informatica and designed data conversions from wide variety of source systems.
- Experienced in SQL Server Reporting Services (SSRS), SQL Server Integration Services (SSIS), Data Transform Services (DTS) and SQL Server Analysis Services (SSAS).
- Experienced in integration of various data sources like Teradata, SQL Server, Oracle, DB2, Netezza and Flat Files.
- Extensively created mappings in Talend using tMap, tJoin, tReplicate, tParallelize, tJava, tJavarow, tDie, tAggregateRow, tWarn, tLogCatcher,tMysqlScd, tFilter, tGlobalmap etc.
- Experienced in working with Horton works distribution of Hadoop, HDFS, MapReduce, Hive, SparkSQL, Impala, Sqoop, Flume, Pig, HBase, Cassandra and MongoDB
- Expert in the field of design, development, and implementation of processes for Data Warehousing and Data Integration projects using Informatica Power Center and Power Exchange.
- Experienced in Data Warehouse/Data mart, OLTP and OLAP implementations teamed with project scope, Analysis, requirements gathering, data modeling, Effort Estimation, ETL Design, development, System testing, Implementation and production support.
- Familiar in using cloud components and connectors to make API calls for accessing data from cloud storage (Google Drive, Salesforce, Confidential S3, DropBox) in Talend Open Studio.
- Solid experience in designing ETL Jobs using Talend Open Studio (TOS) (6.x) and Talend Integration suite (6.x).
- Developed Complex mappings from varied transformation logics like Unconnected/Connected lookups, Router, Filter, Expression, Aggregator, Joiner, Union, Update Strategy and more. Used debugger to test and fix mapping.
- Hands on experience in configuring HDFS and Hadoop ecosystem components like HBase, Solr, Hive, Tez, Sqoop, Pig, Flume, Oozie, Zookeeper and etc.
- Extensive experience in Teradata utilities like MLOAD, FLOAD, TPUMP, FASTEXPORT and TPT for improving target loading performance and have also created complex BTEQ scripts.
- Demonstrated expertise in utilizing ETL tools Informatica power center and Power Exchange 9.x/8.x for developing the Data warehouse loads as per client requirement.
- Design & implement data quality & business rules using IDQ and Power center workflows, Mappings, Mapplets, Exception tables, ad-hoc reporting, Data quality score cards.
- Extensive experience in writing UNIX shell scripts and automation of the ETL processes using UNIX shell scripting, and also used Netezza Utilities to load and execute SQL scripts using UNIX.
- Proficient in Data warehouse ETL activities using SQL, PL/SQL, PROC, SQLLOADER, C, UNIX scripting, Python scripting and Perl scripting.
- Strong knowledge on Data Warehousing Concepts like Ralph Kimball and Bill-Inmon Methodologies
- Excellent experience working on Slowly Changing Dimensions (SCD) Type 1, 2 and 3 to keep track of historical data.
- Strong knowledge of Star Schema, Snow Flake Schema, Fact Table, Dimension Table, Slowly Changing Dimensions, Logical Data Modeling, Physical Modeling and Dimension Data Modeling using Erwin.
TECHNICAL SKILLS
Bigdata and Cloud Techs: Hadoop Framework (Hive, HDFS, Sqoop, Oozie), AWS S3, AWS Redshift, AWS Glue and Airflow.
Databases: Oracle /10g/11g, 12c, SQL Server 2008/2012/2014 , Teradata, Netezza, DB2, MS Access, PostgreSQL, MongoDB and Cassandra
Development Tools: SQL Navigators, Toad, Stylus Studio (XML), XMLDB, Trans-SQL
ETL Tools: Talend 14/15, TIS, TOS, Informatica Power Center 8.x/9.x, SSIS.
Other Tools: XML Spy, Stylus studio, Visual Source Safe. Developer, Oracle Express, Teradata SQL Assistant, Netezza Aginity.
Languages: SQL, PL/SQL, Python
Operating System: UNIX (Sun Solaris), Windows7/8/XP
DB Tools: SQL*Plus, SQL*Loader, Export/Import, TOAD, SQL Navigator, SQL Trace, MLOAD, FLOAD, FEXPORT, TPUMP
PROFESSIONAL EXPERIENCE
Confidential, Chicago, IL
Sr. ETL/Informatica Developer
Responsibilities:
- Prepared High Level Design and Low Level Design based on Functional and Business requirements document of the project and Creating and maintained, detailed support documentation for all ETL processes, developed solutions, including detailed flow designs and drafts.
- Responsible for full data loads from production to AWS Redshift staging environment and complete Data loading from Postgresql to AWS Redshift Data Lake.
- Creating the ETL mappings using various Informatica transformations: Source qualifier, Data Quality, Lookup, Expression, Filter, Router, Sorter, Aggregator etc.
- Imported data from RDBMS environment into HDFS using Sqoop for report generation and visualization purpose using Tableau and worked in Loading and transforming large sets of structured, semi structured and unstructured data.
- Developed processes on both Teradata and Oracle using shell scripting and RDBMS utilities such as Multi Load, Fast Load, Fast Export, BTEQ (Teradata) and SQL*Plus, SQL*Loader (Oracle).
- Worked with the DW architect to prepare the ETL design document and developed transformation logic to cleanse the source data of inconsistencies during the source to stage loading
- Design & Develop ETL workflow using Oozie for business requirements, which includes automating the extraction of data from MySQL database into HDFS using Sqoop scripts.
- Provided extensive Production Support for Data Warehouse for internal and external data flows to Netezza, Oracle DBMS from ETL servers via remote servers.
- Designed and developed an entire DataMart from scratch and designed, developed and automated the Monthly and weekly refresh of the Datamart and Developed several complex Mappings, Mapplets and Reusable Transformations to facilitate one time, Weekly, Monthly and daily loading of Data.
- Used Informatica Power Center for extraction, loading and transformation (ETL) of data in the data warehouse and developed Informatica mappings, transformation, reusable objects by using mapping designer, and transformation developer and Mapplet designer in Informatica Power Center.
- Worked with existing Python Scripts, and made additions to the Python script to load data from CMS files to Staging Database and to ODS.
- Handled Informatica administration work like migrating the code using Export/Import & Informatica Deployment groups, creation of users, creating folders, Worked on Shortcuts across shared and non-shared folders and created various Oracle database SQL, PL/SQL objects like Indexes, stored procedures, views and functions for Data Import/Export.
- Involved in writing the Test Cases and also assisted the users in performing UAT and extensively used UNIX shell scripts to create the parameter files dynamically and scheduling jobs using TWS scheduler.
- Responsible for determining the bottlenecks and fixing the bottlenecks with performance tuning using Netezza Database.
- Used Informatica Power Center Workflow manager to create sessions, batches to run with the logic embedded in the mappings and created complex mappings in Power Center Designer using Aggregate, Expression, Filter, Sequence Generator, Lookup, Joiner and Stored procedure transformations.
- Consolidated data from different systems to load Constellation Planning Data Warehouse using ODI interfaces and procedures
- Developed mappings to load Fact and Dimension tables, SCD Type 1 and SCD Type 2 dimensions and Incremental loading and unit tested the mappings
- Attended SCRUM meetings regularly to discuss the day-to-day progress of the individual teams and the overall project.
- Tuned OBIEE Reports and designed Query Caching and Data Caching for Performance Gains and created integration services, repository services and migrated the repository objects.
- Used heterogeneous data sources XML Files and Flat Files as source also imported stored procedures from Oracle for transformations.
Environment: Informatica Power Center 9.6.1(Repository Manger, Designer, Workflow Monitor, Workflow Manager), UNIX Shell Scripting, Oracle 12c Teradata, Flat files, AWS Redshift, AWS S3, AWS Glue, Python, Hadoop, HDFS, Sqoop, MapReduce, Hive, NoSQL(MongoDB, Cassandra), PostgreSQL, DB2, SQL, Erwin, SQL, PL/SQL, T-SQL, SSRS, NetezzaAginity, Teradata SQL Assistant, SSIS.
Confidential, Seattle, WA
Sr. ETL Developer
Responsibilities:
- Analyze, design, develop, test, implement and troubleshoot integrations between mission critical business applications including cloud based data warehouses and worked closely with Business analysts and Data architects to understand and analyze the user requirements.
- Design and Develop ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
- Used Teradata utilities (TPT, BTEQ) to load data from source to target table and created various kinds of indexes for performance enhancement.
- Installation of Talend Open Studio (TOS) and Configuration along with Java JRE & JDK and performance tuned and optimized various complex SQL queries and transform data to various sources using SQL Server Integration Service and Talend Open Studio.
- Data Extraction, aggregations and consolidation of Adobe data within AWS Glue using PySpark and Create external tables with partitions using Hive, AWS Athena and Redshift
- Created ETL Mapping with Talend Integration Suite to pull data from Source, apply transformations, and load data into target database.
- Delivered MDM stewardship and data governance program. Data monitoring & Notification processes are designed.
- Involved in building the ETL architecture and Source to Target mapping to load data into Data warehouse and developed mappings /Transformation/Joblets and designed ETL Jobs/Packages using Talend Integration Suite (TIS) in Talend.
- Used Teradata external loading utilities like Multi Load, TPUMP, Fast Load and Fast Export to extract from and load effectively into Teradata database
- Created mappings using the transformations like Source Qualifier, XML Source Qualifier, Aggregator, Expression, Lookup, Router, Normalizer, Filter, Update strategy and Joiner transformations.
- Created ETL jobs to load Twitter JSON data into MongoDB and jobs to load data from MongoDB into Data warehouse.
- Used Talendjoblet and various commonly used Talend transformations components like tMap, tDie, tConvertType, tFlowMeter, tLogCatcher, tRowGenerator, tSetGlobalVar, tHashInput&tHashOutput and many more.
- Created complex SCD type 1 & type 2 mappings using dynamic lookup, Joiner, Router, Union, Expression and Update Strategy Transformations.
- Utilized SDLC and Agile methodologies such as SCRUM and designed External and Managed tables in Hive and processed data to the HDFS using Sqoop and Create user defined functions UDF in Redshift.
- Wrote Hive and Pig scripts as ETL tool to do transformations, event joins, filter both traffic and some pre-aggregations before storing into the HDFS.
- Worked on Talend Administration Console (TAC) for scheduling jobs and adding users and used Talend components tMap, tDie, tConvertType, tFlowMeter, tLogCatcher, tRowGenerator and responsible for Performance Tuning at Talend level and SQL queries level.
- In order to increase the performance balanced the input files of slice count against large files and loaded into AWS-S3 Refine Bucket and by using copy command achieved the micro-batch load into the Confidential Redshift.
- Loaded the data from different sources into relational tables with Talend ETL and developed complex Talend jobs mappings to load the data from various sources using different components.
- Wrote SQL overrides in source Qualifier in order to filter the data more effectively at the source level and wrote Complex SQL's in Teradata.
- Explore prebuilt ETL metadata, mappings and DAC metadata and Develop and maintain SQL code as needed for SQL Server database.
- Designed and developed Big Data analytics platform for processing customer viewing preferences and social media comments using Java, Hadoop, Hive and Pig and provide ETL solution to the requirement using BIG DATA Hadoop and imported and exported data into HDFS using Sqoop and Kafka.
- Monitored and supported the Talend jobs scheduled through Talend Admin Center (TAC) and writing above jobs and submitting jobs to do the Teradata restores for migration from Legacy to Teradata.
Environment: Talend Studio, Oracle 12c, XML files, MongoDB, Flat files,, JSON, AWS S3, AWS Redshift, AWS Glue, AWS Athena, HDFS, Hive, HBase, Talend Administrator Console, Agile Methodology, Tableau, SSRS, SQL, PL/SQL, Teradata, Netezza, Aginity, SQL Assistant, Tableau, UNIX Shell Scripting, SQL, T-SQL.
Confidential, Pittsburgh, PA
Sr. ETL Developer
Responsibilities:
- Designed and developed ETL transformations/code for new Data Warehouse Subject areas using SQL or ETL tool, TALEND and interact with Solution Architects and Business Analysts to gather requirements and update Solution Architect Document.
- Performed analysis, design, development, Testing and deployment for Ingestion, Integration, provisioning using Agile Methodology and attended Daily Scrum meetings to provide update on the progress of the user stories Rally and to the Scrum Master and also to notify blocker and dependency if any.
- Imported data from RDBMS environment into HDFS using Sqoop for report generation and visualization purpose using Tableau.
- Created complex mappings in Talend using tMap, tJoin, tReplicate, tParallelize, tJava, tJavarow, tJavaFlex, tAggregateRow, tDie, tWarn, tLogCatcher, etc.
- Designing & Creating ETL Jobs through Talend to load huge volumes of data into Cassandra, Hadoop Ecosystem and relational databases.
- Created Joblets in Talend for the processes which can be used in most of the jobs in a project like to Start job and Commit job and created Context Variables and Groups to run Talend jobs against different environments.
- Worked on Master Data Management (MDM), Hub Development, extract, transform, cleansing, loading the data onto the staging and base object tables.
- Developed processes on both Teradata and Oracle using shell scripting and RDBMS utilities such as Multi Load, Fast Load, Fast Export, BTEQ (Teradata) and SQL*Plus, SQL*Loader (Oracle).
- Loading data from various data sources and legacy systems into Teradata production and development warehouse using BTEQ, FASTEXPORT, MULTI LOAD, and FASTLOAD.
- Developed Talend jobs to populate the claims data to data warehouse - star schema.
- Implemented Change Data Capture technology in Talend in order to load deltas to a Data Warehouse.
- Responsible for improving the Business Objects query performance by tuning the Queries, Database and server parameters.
- Involved in Extensive Design and Code reviews of ETL and Query written in the Teradata environment and developed error logging module to capture both system errors and logical errors that contains Email notification and also moving files to error directories.
- Extensively used Aginity Netezza work bench to perform various DML, DDL etc operations on Netezza database.
- Developed multiple MapReduce jobs in java for Data Cleaning and pre-processing analyzing data in PIG and provided NoSQL solutions in MongoDB, Cassandra for data extraction and storing huge amount of data.
- Utilized Power Query in Power BI to Pivot and Un-pivot the data model for data cleansing and data massaging.
- Responsible for developing, support and maintenance for the ETL (Extract, Transform and Load) processes using Talend Integration Suite and worked in improving performance of the Talend jobs.
- Introduced Tableau Visualization to Hadoop to produce reports for Business and BI team and worked for ETL job design as per criteria in ODI and loaded data table to Teradata server.
- Implemented slowly changing dimensions (SCD) for some of the Tables as per user requirement and performed unit testing and also integration testing after the development and got the code reviewed.
Environment: Talend Studio, Oracle 11i, XML files, Hadoop, Hive, Pig, HBase, MongoDB, Flat files, HL7 files, JSON, TWS, HDFS, Hive 0.13, HBase 0.94.21, MDM, PostgreSQL, Talend Administrator Console, IMS, Agile Methodology, Tableau, Teradata, Netezza, Oracle 11g, Java, SQL, T-SQL, PL/SQL.
Confidential, New York, NY
Sr. ETL Developer
Responsibilities:
- Involved in Extraction, Transformation and Loading of data across different platforms including Hadoop, Big Data database.
- Wrote PL/SQL stored procedures and triggers, cursors for implementing business rules and transformations. Created complex T-SQL queries and functions and provided support to develop the entire warehouse architecture and planned the ETL process.
- Extracted data from flat files, XML files and Oracle, applied business logic to load them in the central Oracle database.
- Migrated ETL jobs to Pig scripts to do Transformations, even joins and some pre-aggregations before storing the data to HDFS and data structure used by NoSQL databases are different from those used by default in relational databases, making some operations faster in NoSQL
- Developed and maintained ETL (Extract, Transformation and Loading) mappings to extract the data from multiple source systems like Oracle, SQLserver and Flat files and loaded into Oracle.
- Performance tuned various mappings, Sources, Targets and transformations by optimizing caches for lookup, joiner, rank, aggregator, sorter transformation and tuned performance of Informatica session for data files by increasing buffer block size, data cache size, sequence buffer length and used optimized target based commit interval and Pipeline partitioning to speed up mapping execution time
- Imported Relational Data base data using Sqoop into Hive Dynamic partition tables using staging tables and imported data using Sqoop from Teradata using Teradata connector.
- Populate or refresh Teradata tables using Fast load, Multi load & Fast export utilities for user Acceptance testing and wrote SQL queries and PL/SQL procedures to perform database operations according to business requirements.
- Created some exclusive mappings in Informatica Developer to load the data from external sources to landing tables of MDM hub.
- Working on a MapR Hadoop platform to implement Bigdata solutions using Hive, Mapreduce, shell scripting, and java technologies.
- Developed mappings/reusable objects/transformations/mapplets by using mapping designer, transformation developer and mapplet designer in InformaticaPower Center
- Extensively used Netezza utilities like NZLOAD and NZSQL and loaded data directly from Oracle to Netezza without any intermediate files.
- Implemented slowly changing dimension to maintain current information and history information in dimension tables.
- Primary activities include data analysis identifying and implementing data quality rules in IDQ and finally linking rules to power center ETL process and delivery to other data consumers.
- Designed and Developed ETL strategy to populate the Data Warehouse from various source systems such as Oracle, Teradata, Netezza, Flat files, XML, SQL Server
- Responsibilities included designing and developing complex mappings using Informatica power center and Informatica developer (IDQ) and extensively worked on Address validator transformation in Informatica developer (IDQ).
- Design & Develop ETL workflow using Oozie for business requirements, which includes automating the extraction of data from MySQL database into HDFS using Sqoop scripts.
- Generated ad-hoc reports in Excel Power Pivot and sheared them using Power BI to the decision makers for strategic planning.
- Created Jobs and Job streams in Autosys scheduling tool to schedule Informatica, SQL script and shell script jobs
- Implemented Real-Time Change Data Capture (CDC) for SalesForce.com (SFDC) sources using Informatica Power Center and implemented Slowly Changing Dimensions for applying INSERT else UPDATE to Target tables.
- Designed complex mappings in Power Center Designer using Aggregate, Expression, Filter and Sequence Generator, Update Strategy, Union, Lookup, Joiner, XML Source Qualifier and Stored procedure transformations.
- Proposed PL/SQL and UNIX Shell Scripts for scheduling the sessions in Informatica and worked with reporting team using the BI interface Business object on improving the business.
Environment: Informatica Power Center 9.3 (Power Center Designer, Teradata, workflow manager, workflow monitor), Oracle 11g, Hadoop, Hive, Pig, MapReduce, HBase, Sqoop, Oozie, IDQ, SQL Server 2010, MDM, TERADATA, PL/SQL, TOAD, Informatica Scheduler, Netezza, TeradataSQL Assistnace, SQL, SSRS, UNIX, Shell Scripting, Autosys, Informatica IDQ, SAP, T-SQL
Confidential
Sr. ETL Developer
Responsibilities:
- Involved in Requirement analysis in support of Data Warehousing efforts and Maintain Data Flow Diagrams (DFD's) and ETL Technical Specs or lower level design documents for all the source applications
- Used SSIS to create ETL packages to validate, extract, transform and load data to data warehouse databases and data mart databases.
- Involved in gathering, analyzing, documenting business requirements, functional requirements and data specifications for Business Objects Universes and Reports.
- Designed and developed Power BI graphical and visualization solutions with business requirement documents and plans for creating interactive dashboards.
- Worked with source databases like Oracle, SQLServer, Teradata, Netezza and FlatFiles.
- Extensively worked with various Active transformations like Filter, Sorter, Aggregator, Router and Joiner transformations
- Developed SSIS Templates which can be used to develop SSIS Packages in such a way that they can be dynamically deployed into Dev, Test and Production Environments.
- Extensively worked with various Passive transformations like Expression, Lookup, Sequence Generator, Mapplet Input and Mapplet Output transformations
- Created complex mappings using Unconnected and Connected lookup Transformations and implemented Slowly changing dimension Type 1 and Type 2 for change data capture
- Worked with various look up cache like Dynamic Cache, Static Cache, Persistent Cache, Re Cache from database and Shared Cache
- Responsible for the performance tuning of the ETL process at source level, target level, mapping level and session level and Worked with various SSIS objects like Mappings, transformations, Mapplet, Workflows and Session Tasks
- Involved in writing Teradata SQL bulk programs and in Performance tuning activities for TeradataSQL statements using Teradata EXPLAIN
- Extensively used debugger to test the logic implemented in the mappings and performed error handing using session logs.
- Design, developed and Unit tested SQL views using Teradata SQL to load data from source to target.
- Writing Complex T-SQL Queries, Sub queries, Co-related sub queries and Dynamic SQL queries and performance tuning of Stored Procedures and T-SQL Queries.
Environment: SQL Server Integration Service (SSIS), Oracle, Power BI, SQL Server, Sybase, SQL, PL/SQL, Cognos, Unix AIX and windows XP, Netezza, Teradata, IDQ, Business Objects, MDM, T-SQL, SSRS.
