Data Science Analyst Resume
Louisville, KY
SUMMARY:
- Over 5 years of Industry experience as aDataAnalyst and Data Science Analyst with solid understanding ofOptimization, Simulation, Statistics, Data mining, Machine learning, neural nets and Forecasting techniques .
- Experienced in supervised learning including classification/regression and regularization algorithms, unsupervised learning such as topic models
- Analytically identify, conduct and track experiments (A/B testing) for channel migration initiative to change customer behavior.
- Strong experience of developing and maintaining business reports including ad hoc reporting, trending and analysis
- Expertise in using T - SQL for developing complex Stored Procedures, Triggers, Tables, Views, User Functions, User profiles, Relational Database Models andDataIntegrity, SQL joins and Query Writing.
- Expertise in writing SQL Queries, Dynamic-queries, sub-queries and complex joins for generating Complex Stored Procedures, Triggers, User-defined Functions, Views and Cursors.
- Used R to perform text mining of social media data for sentiment analysis and topic classification.
- Forecasted customer behavior for various companies to increase customer sales, retention, or engagement.
- Written several shell scripts using UNIX Korn shell for file transfers, error logging, data archiving, checking the log files and cleanup process.
- Experience in creating UNIX scripts for file transfer and file manipulation.
- Worked on data modeling and produced data mapping and data definition specification documentation.
- Excellent understanding and working experience of industry standard methodologies like System Development Life Cycle (SDLC), AGILE Methodologies.
- Worked on data warehousing, ETL, SQL, scripting.
- Database management fundamentals including creating/updating/dropping tables, developing stored procedures.
- Strong experience with data mining and management with huge data sources
- Proficient in Gathering Requirements, Business Case Development to Rollout, Production and Maintenance with a solid understanding of Business Processes and analyzing them and documenting.
- Excellent experience in writing SQL queries to validatedatamovement between different layers in datawarehouse environment.
- Extensive ETL testing experience using Informatica 9x/8x and Datastage.
TECHNICAL SKILLS:
Programming Languages: C++, Python, Java, PHP, JSP, SQL/PLSQL, R Programming
Data Science: Regression, Classification and Clustering models, Neural Networks
Tools/IDE: Ipyhton Notebook, Scikit-Learn, NLTK, Eclipse, Git, Weka, Pandas and Numpy
DataModeling: Star-Schema Modeling, Snowflake-Schema Modeling, FACT and dimension tables, Pivot Tables.
Databases: Oracle11g, DB2, SQL Server, MySQL, MS- Access, Flat Files, XML files.
Operating Systems: Windows, Linux and UNIX
Scheduling Tools: Autosys, Maestro (Tivoli)
ETL/Datawarehouse Tools: Informatica, SAP Business Objects, Web Intelligence, Tableau, IBM Datastage.
PROFESSIONAL EXPERIENCE:
Confidential, Louisville KY
Data Science Analyst
Responsibilities:
- Designed and implemented Naive Bayes classifier algorithm to classify customers with prior information in training sets
- Created and maintained Database Objects (Tables, Views, Indexes, Partitions, Synonyms, Database triggers, Stored Procedures) in the data model.
- Performed 5-fold cross-validation to search for best tunable parameter with low false positives and high true positives.
- Involved with Data Analysis Primarily Identifying Data Sets, Source Data, Source Meta Data, Data Definitions and Data Formats.
- Cleaned text contents (e.g., removed punctuation, tokenized the article contents into words, filtered out stop words) with natural languages processing package NLTK in Python
- Enforced referential integrity in the OLTP data model for consistent relationship between tables and efficient database design.
- Create insightful and engaging data visualizations, ranging from customer level demographics to financial impact to geospatial patterns
- Craft presentations to engage stakeholders and drive strategy.
- Walked through the LogicalDataModels of all source systems fordataquality analysis.
- Handled performance requirements for databases in OLTP and OLAP models.
- Worked on coming up with a statistical model to calculate the customer risk score of the customers.
- Created advanced supervised and unsupervised machine learning algorithms specifically for tasks with severe class imbalance, such as fraud detection, fault detection, intrusion detection.
- Involved in the integration ofdatacoming from different sources.
- Involved in the creation, maintenance ofDataWarehouse and repositories containing Metadata.
- Designed different type of STAR schemas like detaileddatamarts and Plandatamarts, Monthly Summarydatamarts using ER studio with various Dimensions Like Time, Services, Customers and various FACT Tables.
- Developed and maintaineddatadictionary to create metadata reports for technical and business purpose.
Environment: PL/SQL, Business Objects XIR2, ETL Tools Informatica9.5/8.6/9.1 Oracle 11G, Teradata V2R13/R14.10, Python, Pandas, Weka, R Studio
Confidential, NYC, NY
Data Science Analyst, Insurance
Responsibilities:
- Determineddatarules and conducted Logical and Physical design reviews with business analysts, developers and DBAs.
- Invented novel feature selection methods to identify features that can be shared across different tasks. Application areas include sentiment prediction.
- Applied benchmark machine learning algorithms, such as SVMs, Logistic Regression and HMMs to numerous tasks in natural language processing.
- Worked ondatamapping process from source system to target system.
- Developed a two-way ANOVA model which helps in estimating the market value of the properties based on their location and size
- Built this ANOVA model in R, imported multcomp package in R for multiple comparison and performed diagnosis analysis for homoscedasticity using Levene's test.
- Worked on Performance Tuning of the database which includes indexes, optimizing SQL Statements.
- Implemented a supervised learning model to predict whether a loan will be defaulted or not
- Trained the model using Logistic Regression, Naive Bayes and Decision trees. Performed cross validation to tune the parameters and improve the test accuracy of the models.
- Created tables, views, sequences, triggers, tablespaces, constraints and generated DDL scripts for physical implementation.
- Worked at conceptual/logical/physical data model level using Erwin according to requirements.
- Performed data mining on data using very complex SQL queries and discovered pattern.
- Used SQL for Querying the database in UNIX environment.
- Involved in extensiveDatavalidation by writing several complex SQL queries and Involved in back-end testing and worked withdataquality issues.
- Performed data analysis, Data Migration and data profiling using complex SQL on various sources systems including Oracle
- Designed semantic layer data model. Conducted performance optimization for BI infrastructure.
Environment: PL/SQL, Business Objects XIR2, ETL Tools Informatica9.5/8.6/9.1 Oracle 11G, Teradata V2R13/R14.10, Python, Pandas, Weka, R Studio
Confidential
Data Analyst / Data Modeler
Responsibilities:
- PerformedDataAnalysis andDataProfiling and worked ondatatransformations anddataquality rules.
- Designed both 3NFdatamodels for ODS, OLTP systems and dimensionaldatamodels using star and snow flake Schemas
- Wrote ad-hoc SQL queries and worked with SQL and Netezza databases
- Wrote complex SQL queries for validating thedataagainst different kinds of reports generated by Business Objects.
- Performed in depth analysis indata& prepared weekly, biweekly, monthly reports by using SQL, MsExcel, MsAccess, and UNIX.
- Extensively used Tableau to create interactive dashboard & MySQL and Excel for analysis and finding the answer for business questions.
- Document variousDataQuality mapping document, audit and security compliance adherence.
- Extensively worked Data Governance, i.e. Metadata management, Master data Management, Data Quality, Data Security
- Possessed strong Documentation skills and knowledge sharing among Team, conducteddata modeling review sessions for different user groups, participated in sessions to identify requirement feasibility.
- Worked with data investigation, discovery and mapping tools to scan every single data record from many sources.
- Performed data analysis and data profiling using complex SQL on various sources systems including Oracle, SQL server and DB2.
- Performed data analysis and data profiling using complex SQL on various sources systems including Oracle 10g/11g and Teradata.
- Interactive data visualization using Tableau focused on business intelligence data
Environment: ER/Erwin, PL/SQL, Business Objects XIR3, ETL Tools Informatica 9.5/8.6/9.1 Oracle 11G, Teradata V2R12/R13, SQL, TOAD, PL/SQL, Flat Files, Tableau
Confidential
Data Analyst / Data Modeler
Responsibilities:
- Extensive hands-on technical knowledge of Data Warehouse Architecture, ETL, Database and SQL performance tuning
- Experienced inDataTransformation andDataMapping from source to target database schemas and alsodatacleansing.
- Created various PhysicalDataModels based on discussions with DBAs and ETL developers.
- Worked ondatamapping process from source system to target system.
- Conducted Gap Analysis, Risk Analysis to identify and develop Risk Mitigation strategies.
- Extensively used Star and Snowflake Schema methodologies.
- Developed and maintainedDataDictionary to create Metadata Reports for technical and business purpose.
- Worked on Performance Tuning of the database which includes indexes, optimizing SQL Statements.
- Designed Universes and Reporting using SAP BO, including Report Design using Desktop Intelligence
- Created tables, views, sequences, triggers, table spaces, constraints and generated DDL scripts for physical implementation.
- Worked at conceptual/logical/physical data model level using Erwin according to requirements.
- Performed data mining on data using very complex SQL queries and discovered pattern.
- Used SQL for Querying the database in UNIX environment.
- Involved in extensiveDatavalidation by writing several complex SQL queries and Involved in back-end testing and worked withdataquality issues.
- Performed data analysis, Data Migration and data profiling using complex SQL on various sources systems including Oracle.
- Designed semantic layer data model. Conduct performance optimization for BI infrastructure.
Environment: DB2, CA Erwin 7.0, Oracle 11g, VISIO, MS-Office, SQL Architect, TOAD Benchmark Factory, SQL Loader, PL/SQL, SharePoint, INFORMATICA, VISIO, Oracle, MS-Excel
