Data Architect Resume
Charlotte, NC
SUMMARY:
- This page is a condensed version of my professional experience. Individual project details are listed on page two and onwards in reverse - chronological order (latest first).
- Data Architect with Data Science chops, Data Modeling experience and ETL background.
- Fluent with Cloud and On-premise database implementations.
- 13 years of IT experience with first 2 in a PC sales startup during undergraduate course.
- Enthusiastic about data engineering projects. Competent independently and in teams.
- Avid runner, curious, amenable.
TECHNICAL SKILLS:
- Data modeling
- Statistical Data Mining (Data Science)
- Data Lineage
- Data Profiling
- Data Warehousing
- Toolkit
- SQL
- Python (2.x, 3)
- ERwin (7.x, 9.x)
- Datastage ETL (8.x, 9.x, 11.x)
- Informatica ETL
- Infosphere Metadata Workbench
- Unix
- Oracle 9i, 10g, 11g
- DB2
- Netezza
- Tableau
- Redshift
PROFESSIONAL EXPERIENCE:
Confidential, Charlotte, NC
Data Architect
Responsibilities:
- Design a new data model for Confidential
- Data profiling on existing systems of record
- Data modeling for application workflow and UI navigation
- Enhance data model for new and incoming feeds
- Data governance through data quality, metadata management, delegation to data stewards.
Technology: ERwin 9.x , MS SQL Server 2014
Confidential, Charlotte, NC
ETL Senior Developer
Responsibilities:
- Discuss data model changes for new requirements.
- Data Mining for business analysts.
- Design and develop ETL jobs.
- Some achievements:
- Data model changes to add business rules for an administrative schema (classification value tables) saving a lot of procedural code writing effort.
- Data model changes for small to medium complexity ETL projects.
- Some interesting ETL work done at Confidential:
- Converted a 45-hour batch process that ran on Oracle’s Endeca server to a new Datastage batch that consists of about 35 new jobs. The resulting process runs in 2.5 hours. Data is integrated from a multitude of databases and flat files. It produces extract files as output up to the size of 200 GB.
- Implemented and tuned a complex set of custom aggregation transformations in Datastage.
- Performance tuned ETL jobs involving heavy usage of LONGVARCHAR fields that ended up using a lot of Datastage scratch space causing jobs to abort.
Technology: ERwin , Infosphere Datastage, GNU/Linux, Oracle 11g, AutoSys, StarTeam
Confidential, Philade lphia, PA
Lead ETL Developer
Responsibilities:
- Implemented data model changes for ordering system by separting NMWA (Notify Me When Available) from the Cart events table.
- Understand existing data warehouse architecture and identify pain points and bottlenecks for performance improvement. Some of these improvements were just result of new tool and hardware, some due to redesign of ETL processes.
- POC work for data lineage using IBM Metadata Workbench
- Design high-volume XML-based data integration jobs in the new IBM Infosphere suite.
- Datastage job development to read XML data, to use Netezza partitioned reads/TWT loads (using dataslice id), Oracle (table partitioning/mod partitioning) etc.
- Review/guidance for Offshore team.
Technology: ERwin, Infosphere Datastage, Metadata Workbench, GNU/Linux, Netezza, Altova XML Spy, AutoSys, SVN/Tortoise
Confidential, Charlotte NC; Dallas TX
ETL Developer
Responsibilities:
- Developed some of the core ETL jobs for Confidential .
- Interesting work done on Datastage:
- Custom aggregation/summary (nested if else logic) using KeyChange variable, and Transformer Stage variable manipulation. Summarizes from ~10mil records to a few hundred thousand records and loads to a summary table.
- Performance tuned several jobs that overloaded the 4-node Oracle DB by redesigning jobs to take advantage of 32-node ETL server.
- Developed a job for file validation - count the number of records against count in Trailer record. This requires careful handling of datastage parallelism vs record sequencing, use of transformer for link sorting, stage variables and constraints, Otherwise/Log and AbortAfterRows properties.
- Debugged and resolved a bug with XML Output stage that involves handling of APT STRING PADCHAR using Convert function.
- Use of awk script in a pre-job routine to handle multiple files of a pattern with interleaved multiple header/trailer in each file.
Technology: ERwin, Infosphere Datastage, GNU/Linux, Oracle 10g, AutoSys, StarTeam, QualityCenter.
Confidential, Dallas, TX
ETL Developer
Responsibilities:
- Data extraction project for Basel compliance. This was a fairly complex work to do on SQL and was done on time and on budget using DB2 SQL and AIX.
- Design of horizontal and vertical splicing of a very large table (200 million records, 300 columns) for efficient access, security and performance.
- Data modeling through “reverse engineering ” the amalgam of data after an ac quisition.
- ETL using Datastage, SAS, Informatica (and sometimes SQL) for number of projects - Consumer IVR, Privacy and Do Not Disturb etc.
- Performance tuned a few ETL/purge processes:
- SQL/SAS: Redesigned the core housekeeping script for purging on warehouse. Purge script would loop through records iteratively on different DB2 nodes, but was found to have leaving certain database nodes underutilized increasing the transaction volume between commit points and thus filling up db logs. The redesigned solution would loop through records based on record count instead of node number. This helped pumping small manageable counts of records to all nodes equally.
- Informatica: Modified the Source Qualifier of a mapping (job) to 8-way partitioned reader based on mod value of an evenly distributed integer field. Split the mapping into two parts to work around non-availability of DB2 coordinator node on partitioned loads. Enabled bulk load to load to DB2 table. The Lookup Transformation present in the mapping was made a Persistent Lookup and the Lookup Cache was re-directed to a Unix mount point with large free space. The tuned process saved 5.5 hours of running time reducing the average run time to about 20 minutes.
- SAS: Tuned a slow running SAS extraction process by splitting up data fetch process in proc sql to multiple partitions, removed redundant sorts and combined redundant proc sqls, eliminated a recursive update.
Technology: ERwin, AIX, Infosphere Datastage, SQL, Base SAS, DB2 UDB, Oracle 10g, TOAD 8.6, UNIX Shell scripting, Tivoli M aestro Scheduler
Confidential, Irving
VB Developer/DB support
Responsibilities:
- Analyze VB code for issues/tickets and provide resolution within SLA.
- Involve in DB support activities and work with DBAs and developers in Dev Database regions
- Maintenance of Unix scripts that helped migration and clean up during Production Migration of data as well as Dev Region cycling
- Maintenance of ClearQuest reports on ClearQuest’s Unix host.
Technology: AIX, SQL, Oracle, ClearQuest, ClearCase
