Data Science Consultant Resume
Portland, OR
SUMMARY:
- An innovative Data Scientist with over 7 years of corporate experience. Accomplished background in Machine
- Learning (Supervised, Unsupervised), Web Scraping, Data Acquisition, Data Loading, Exploratory Data
- Analysis, Data Mining, Data Modeling, Business Modeling, Requirement Gathering, Technical Documentation, and strong domain expertise in Financial & Health Care claims management technologies involved in including Hadoop HDFS, SAP HANA,VB.Net, SOA(web services, XML), SQL and code. Proficient in statistical programming languages like R, SAP HANA(PAL) and Python (Pandas, NumPy, Sckit - Learn, Beautiful Soup) and Mahout for implementing machine learning algorithms in production environment.
TECHNICAL SKILLS:
Platform: Windows 98/2000/XP/Vista/7, UNIX, Mac OSX, LINUX, ERP, SOA.
Databases: HADOOP (HDFS),NOSQL,SAP HANA,IBM DB2,Oracle, MS Access, MS SQL,PIG,HIVE, SPARKSQL
Languages: R, SPSS, Python, SQL,C, HTML, VB scripting, .ASP, Visual Studio, SOA (XML)
Packages: Pandas, NUMPY, SCKIT-Learn, Beautiful SOUP, GGPLOT2, CARET, MAHOUT, DPLYRBusiness Process Modeling Tools: MS VISIO, RUP Tools (Rational Requisite Pro, Rational Rose, RationalClear Case)
Documentation Tools: SharePoint 2013, MS Office(Word/Excel/Power Point/ Visio)
WORK EXPERIENCE:
Data Science Consultant
Confidential - Portland, OR
Responsibilities:
- Performed Data Profiling to learn about user behavior
- Merged user data from multiple data sources
- Performed Exploratory Data Analysis using R and Hive on Hadoop HDFS
- Prototype machine learning algorithm for POC (Proof Of Concept)
- Performed Data Cleaning, features scaling, features engineering Developed novel approach to build machine learning algorithm and implement it in production environment
- In real - time association rules were implemented which uses prior probabilities.
- SAP HANA platform was used implementation which provides several mining algorithms, associated rules etc.
- On the front end SAP Lumira was used.
- By stringing the 1500 datacodes were transformed into 24 different unique features.
- Developed Performance metrics to evaluate Algorithm's performance.
Environment: TERADATA, Oracle, HADOOP (HDFS), RStudio, Python, JAVA, SAP HANA, HIVE, SAP
Principal Data Scientist
Confidential - Chicago, IL
Responsibilities:
- Performed Data Profiling to learn about user behavior.
- Merged user data from multiple data sources.
- Performed Exploratory Data Analysis using Python and Hive on Hadoop HDFS.
- Prototype machine learning algorithm for POC (Proof Of Concept).
- Performed Data Cleaning, features scaling, features engineering Performed ad - hoc data analysis for customer insights using Hive.
- Developed Performance metrics to evaluate Algorithm's performance.
- Used RMSE score, F-SCORE, PRECISION, RECALL, and A/B testing to evaluate recommender's performance in both stimulated environment and real world.
- Fine tune the algorithm using regularization term to overcome the problem of over fitting.
Environment: TERADATA, Oracle, HADOOP (HDFS), PIG, MySQL, RStudio, Python, JAVA, MAHOUT, HIVEPIG, SPARK
Data Scientist
Confidential - Reston, VA
Responsibilities:
- Defined Project Scope, project Charter & Business Case
- Prototype machine learning algorithm for POC (Proof Of Concept)
- Performed Data Cleaning, features scaling, features engineering Developed predictive models for use in machine learning platform using the scikit - learn python framework
- Improved statistical models using learning curves, parameter curves, feature selection, and regularization.
- Performed ad-hoc data analysis for customer insights using SQL using Amazon AWS Hadoop cluster
- Developed Map Reduce pipeline for feature extraction
- Implemented Support Vector Machine (lite)
- Performed Principal Component Analysis (PCA) & Linear Discriminate Analysis(LDA)
- Fine-tuned low bias & High variance trade off Defined the technical requirements of the analytic solutions.
- Defined the data requirements of the analytic solution.
- Worked on commercial data from desperate source systems, built data models and transformed data to provide added value in IT applications by streamlining processes, reducing cost, maximizing profits & rolling ut business solutions that met one of the objectives
- Worked closely with subject-matter experts and business analysts from SAP and non-SAP systems and platforms, and investigating statistical and predictive and prescriptive patterns in the data to build business solutions
- Made iterative changes to analytic/predictive models and decision logic embedded in operational applications and business process platforms
- Worked with multiple relational, dimensional, and OLAP databases
Environment: MS SQL, Oracle, HADOOP (HDFS), PIG, MySQL, SAP Sybase, RStudio, Python, JAVA,.NETHIVE, HADOOP HDFS, PIG, MAHOUT
Data Scientist
Confidential - Salt Lake City, UT
Responsibilities:
- Data analysis and visualization (Python, R,)
- Designed, implemented and automated modeling and analysis procedures on existing and experimentally created data
- Increased pace & confidence of learning algorithm by combining state of the art technology and statistical methods; provided expertise and assistance in integrating advanced analytics into ongoing business processes
- Parsed data, producing concise conclusions from raw data in a clean, well - structured and easily maintainable format
- Implemented Topic Modelling, PASSIVE AGGRESSIVE & other linear classifier models
- Perform tfidf weighting, normalize
- Performed scheduled and adhoc data driven statistical analysis, supporting existing processes
- Developed clustering models for customer segmentation using R
- Created dynamic linear models to perform trend analysis on customer transactional data in R
- Performed Topic modeling
Environment: R, SQL, Python, TABLEU, SAP HANA, SAS, JAVA, PCA & LDA, regression, logistic regressionrandom forest, neural networks, Topic Modeling, NLTK, SVM(Support Vector Machine), JSON, XML, HIVEHADOOP, PIG, MAHOUT
Data Scientist
Confidential - Princeton, NJ
Responsibilities:
- Responsible for predictive analysis of credit scoring to predict whether or not credit extended to a new or an existing applicant will likely result in profit or losses.
- Primarily used R packages for the data mining tasks.
- Participated in all phases of data mining; data collection, data cleaning, developing models, validation and visualization.
- Data for modeling was collected using SQL by querying several tables. The extracted tables were further appended or merged to create tables for modeling using SAS PROC MERGE and PROC SET procedures.
- Adopted principal component analysis, PROC PRINCOMP to reduce the dimension of the data. The missing values were replaced if applicable with the group average using proc means.
- Computed Credit Risk Parameters such as Probability of Default (PD) and Loss Given Default (LGD) and
- Exposure at Default (EAD).
- Used logistic regression (PORC LOGISTIC), clustering (PROC CLUSTER) and multivariate modeling to provide valuable analytical insights.
- Used Kolmogorov - Smirnov test (K-S test or KS test) to measure the quality of the models.
- Used PROC GRAPH for generating various graphs and charts for analyzing the different features.
- Used k-fold cross validation to avoid over fitting.
Environment: .Net, SAS, R, Python, Oracle, IBM DB2, MS SQL, HIVE, HADOOP, PIG, MAHOUT
