Data Scientist Resume
SUMMARY
- Over eight years’ experience solving business problems for Fortune 500 clients - using Machine Learning, Exploratory Data Analysis, and developing data pipelines using Big data technologies & Python data stack in Client facing projects where I got experience working in entire data science life cycle.
TECHNICAL SKILLS
- Machine Learning
- Statistical Modeling
- Predictive Modeling
- Deep Learning
- Natural Language Processing
- Cloud Computing
- Big Data
- Exploratory Data Analysis
- SQL
- Web Scraping
PROFESSIONAL EXPERIENCE
Confidential
Data Scientist
Responsibilities:
- Used Word Embeddings (deep learning) to build a classification model for categories based on job descriptions and extracted various website meta data from Confidential Terabyte database called “docbase” and created Name -> Nickname mappings using Named Entity Extraction.
- Developed visualization dashboards which report insights related to various back office teams such as Agent’s Quality, Job Count trends, Active vs. Expired trends, Job source trends with multi level filters.
- Developed process automation scripts using Python and Selenium which helps reduce the human effort in monotonous activities.
Tools: Python, Pandas, Hive(hadoop) to pull the data, Machine Learning, scikit learn, Deep Learning(word embeddings - gensim), Natural Language Processing Toolkit(NLTK) for entity extraction. Used Confidential ’s internal tools to develop data visualization dashboards.
Confidential
Responsibilities:
- Collaborated with Broadridge’s data engineering team to extract data from their greenplum database.
- Extracted numerical features from the existing variables and performed feature selection using recursive feature elimination.
- Converted the categorical data into one-hot encoded features and removed last binary feature to avoid dummy variable trap.
- Final features used:
- Category(large, micro, mid, small)
- Service type(anonymized)
- Meeting Mail Duration(Mail Date - Meeting Date)
- Shares Per Account (Total Shares / Total Accounts)
- Total voted by user till Mail Date.
- One hot encoded features of Voting type(For, Against, Abstain, Others).
- Count of votes till date for each category.
- Adding CountVectorizer on Policy Rules Name improved the validation score.
- Performed preprocessing - scaling all the features in the same range, standardize the features.
- Evaluated multiple machine learning classification models and ended up using Gradient Boosting with 0.92 AUC score.
- Tested the model using cross validation and evaluated the results on validation subset.
Tools: Python pandas, scikit learn and scipy for data preprocessing, bag of words, and machine learning models, matplotlib/seaborn for exploratory data analysis, XGBoost for final model predictions.
Confidential
Responsibilities:
- Improved the previous model AUC score from 0.78 to 0.89 score.
- Developed data pipeline using Apache Sqoop for data ingestion from Oracle, Hive for data processing.
- Extracted numerical features from the existing variables and performed feature selection using recursive feature elimination.
- After applying preprocessing, feature engineering and feature selection, the following features were extracted:
- One month average call duration( (start time - end time) / total count) computed for inbound, outbound, International, and Roaming.
- Counts of inbound, outbound, international, roaming, and voip calls over six months.
- Account features such as account balance and account type( corporate or consumer).
- Billing features such as Bill Amount, Number of times average spending < 300 rupees in the past 6 months(binary).
- User features such as age, tenure, location, total revenue.
- Quality features like Call drop rate(call drop after answer/calls answered), Mean opinion score(1 to 5).
- Generic features such as handset type.
- Data features such as data usage, page response success rate, average page response delay, page response success rate(total page response success/total page requests).
- Lag features(-1, -2) for Bill Amount, Call usage.
- We evaluated several models on the data and ended up using Gradient Boosting for prediction.
- We improved the model by choosing the better parameters using hyper parameter tuning.
Tools: Python pandas, scikit learn and scipy for data preprocessing, and machine learning models, matplotlib/seaborn for exploratory data analysis, XGBoost for final model predictions. Big Data(Hadoop-Hive for data preprocessing), Apache Sqoop for data ingestion.
Confidential
Developer
Responsibilities:
- Developed an Confidential using meteorjs framework with Mongodb as the database.
- The app has features such as User Signups, Todo lists, Task assignments, file attachments, schedules, user management.
Tools: Javascript, Meteorjs, Mongodb, Twitter bootstrap.
Confidential
Data Specialist
Responsibilities:
- Developed the data pipeline using python, pandas, oracle, sqlite, and tornado, which collects the data from Oracle database, performs quantitative data analysis based on the business specifications using pandas, stores the computed results in sqlite database and uses Tornado and Gramex to render data visualizations.
- Developed a classification model using random forests using sixteen anonymized features to predict four types of risk scores.
- Perform quantitative analysis and statistical analysis based on the business requirements such as Regression(for identifying the positive and negative trend of risk scores over months), and time series analysis(for finding the trend of queries, recruitment, and issues).
- Developed a Recruitment plan comparison dashboard which compares the plan for past and current month, which helped the client identify “plan fraud” committed by trial leaders, which was later reported to Client’s Vigilance department.
Tools: Python, Pandas for quantitative analysis, Machine Learning, scikit learn for Machine Learning, Tornado for developing the web server.
Confidential
Responsibilities:
- Confidential approached us to extract insights from Sales data to improve the efficiency of their sales team across geography.
- Developed a visualization dashboard which included metrics such as WoW%, change, MoM% change, and top 5 & bottom 5 sales regions.
Tools: used: Python, pandas, Gramex(inbuilt data visualization tool).
Confidential
Responsibilities:
- Use Twitter Stream API to continuously monitor a search term as a Proof of Concept to show to clients.
- Project was adopted as a boilerplate across several projects in the company.
- Inputs can be given in a csv file.
- The dashboard shows the visual representation of sentiments.
Tools: used: Python, tweepy for pulling the data from twitter.
Confidential
Engineer
Responsibilities:
- Developed quantitative analyzer using python and tabular which computes the numbers required for report generation for daily,weekly, and monthly reports(monotonous) which are used by over 12 colleagues in the Project Monitoring team for reporting.
- The quantitative analyzer was developed out of my interest to save time & effort, for which I got good appreciation from my Management and the Client.
- I was responsible on reporting for three zones - Raw Material Handling Plant, Outdoor Pipelines, and Blast Furnace.
- Collect data from Site Engineers, segregate the data accordingly and prepare the reports.
Confidential
Data Analyst
Responsibilities:
- Group has businesses across Education(2 colleges), Construction industry, Raw material production, imports and exports.
- Responsible for performing analysis, prepare reports using powerpoint every week and send it to the Management.
- Prepare reports with insights containing the progress and delays.
- Management used to consider action points based on the analysis reported.
- Most of my work was related to the Construction industry where I used to prepare analysis on the sub-contractors.
- Management used to conduct meetings with sub-contractors for tracking the work, and setting targets.
