We provide IT Staff Augmentation Services!

Data Scientist (manager Operations Research & Data Mining) Resume

5.00/5 (Submit Your Rating)

FL

SUMMARY:

  • Overall 15 years of experience as Data Analytics Professional in various service and manufacturing domains, with 3 years of experience in hands on coding and developing predictive models.
  • Machine learning expert with experience in developing predictive models to solve business problems.
  • Business analytics professional who can not only talk to business leaders to understand their requirement but also code.
  • Proficient in using Python, R, SQL, Hadoop ecosystem for extracting data and building predictive models.
  • Very good knowledge in various Statistical Methods like Time Series Analysis, Statistical Testing, Correlation, Multivariate Analysis, Forecasting, Business Intelligence tools and application of Statistical Concepts.
  • Analytics and Business Intelligence leader with a strategic focus on providing data - driven insights and end-to- end solutions to business problems.
  • Proven experience in preparing and presenting performance dash boards to top level executives of organization.
  • Process improvement expert with experience in leading cross functional team to improve productivity and service levels
  • Quick learner of various new technical concepts in machine learning/data science/deep learning field.

TECHNICAL SKILLS:

Analytics Programming & Software: Python, R, Hadoop, Pig, Hive, SAS Enterprise Guide, BASE SAS

Relational & NoSQL Databases: Oracle 11g/12c, SQLServer 2008 R2, MySQL, Mongo DB, Elasticsearch

Big Data skills: Hadoop, AWS EMR, Talend

Web Analytics: A/B Testing, Google analytics

Data visualization: Tableau, d3.js

PROFESSIONAL EXPERIENCE:

Confidential, FL

Data Scientist (Manager Operations Research & Data Mining)

Responsibilities:

  • Whistle post is a sign, which gives indication to train driver to Confidential .
  • Confidential need to maintain inventory of whistle posts for auditing purpose.
  • Now, manual inspectors take inventory by walking along the track.
  • Objective is to identify the whistle posts along the track from the photos taken by train during the journey.
  • Used deep learning model (inception model) of open source project tensorbox to identify the whistle post from the images and achieved 95% accuracy.
  • Confidential from AWS is used for running deep learning model.
  • Now, business decisions are not data driven resulting in unnecessary locomotive movements resulting in wastage equivalent to 2 un-utilized locomotives (each locomotive costs more than $1M) and rework on locomotives.
  • Provide business insights regarding wheel wear rate, wheel machining and wheel replacement patterns to improve the overall wheel life. Each axle with two wheels costs $8000 per unit. Now, business has no insights regarding metrics of performance across different locomotive types.
  • Developed decision tree model to identify a whether locomotive requires machining or not. Achieved a specificity of 70% in predicting machining event, by adjusting threshold probabilities. Now, this model is in testing phase.
  • Analytics regarding variations in wheel life across different types of engines. And, also identified areas for operations improvement to increase the life.
  • Also, tested business hypothesis that 'wheel wear rate is correlated to tonnage and mileage.

Tools: used - Python (Pandas), R (Caret, ctree)

Confidential

Data Scientist

Responsibilities:

  • Wayside sensors collect data related to train axle and wheel movement.
  • Now, this data is not being used for any analytic purposes.
  • Purpose of the project is to leverage this data to do preventive maintenance of locomotive instead of reactive maintenance.
  • Extracted data from Oracle database using ETL tool, TOAD, and wrote SQL queries to combine data related to Locomotive, Truck, Axle and Wheel (12M rows).
  • Exploratory data analysis was done using Python (Pandas).
  • Extracted features from sensor data viz. mean, median, variance of sensor readings.
  • Developed classification model to predict locomotive failures using way side sensor data
  • Applied ensemble models to classify the locomotive failures.
  • Further, to improve the accuracy of the model, applied anomaly detection techniques.

Tool: used - Python (Pandas), R (Caret, randomForest), SQL

Confidential

Data Scientist

Responsibilities:

  • Scan data was extracted from Oracle database using ETL tool, TOAD and wrote SQL queries to prepare data for modeling purposes.
  • Developed regression model to predict wooden tie deterioration. Features for this model include tonnage carried, speed of the train, and curvature of the track.
  • Built an ensemble model with Random Forest, XGBoost and Support Vector Regression model to achieve MAPE value of ~10%.
  • Results of the model are presented on a Tableau dashboard to business
  • Identified and highlighted patterns, data quality issues, and opportunities to business partners and Confidential Technology.

Tool: used - Python (Pandas), R (Caret, randomForest, XGBoost), SQL

Confidential, CA

Data Scientist (Intern)

Responsibilities:

  • Acquired data related YouTube videos by querying channel meter API with Python and stored data in POSTGRES database on AWS remote Linux machine.
  • Used Python’s psycopg2 package to connect to POSTGRES to insert, update and acquire data for descriptive and predictive modeling.
  • Queried freebase API to gather topics related to YouTube videos to generate features for predictive modeling.
  • Generated data frame with 45 features for predictive modeling of number of views for seven days views after a YouTube video is uploaded.
  • Developed descriptive and predictive models using R, Python, POSTGRES. Performed model comparisons for various scenarios.
  • Applied Analytics Techniques viz. Regression, Variable Transformations, Clustering.
  • Applied clustering algorithms (hierarchical clustering with Euclidean distance as a measure on first 7 days views) to cluster YouTube videos to predict views for 8th, 9th and 10th day after uploading the video
  • Applied Linear Regression, Support vector regression and Random forest regression models for predicting number of views and achieved a pseudo R-square value of 0.76

Confidential, CA

Data Scientist (Intern)

Responsibilities:

  • Executed data acquisition, cleaning procedures on structured and non-structured data sources using Python and R in Sourcing Big Data project.
  • Used Hadoop Map-Reduce to identify similar parts across 5M parts in oil & rig designs based on mechanical parts descriptions. Used AWS EMR to demonstrate performance of my map-reduce code. Reduced execution time from 9 hours to 2 hours using map-reduce for clustering the similar parts.
  • Also provided solution with Elasticsearch and Python to group 5M parts.
  • Built a prototype Elasticsearch index with 500,000 parts.
  • Clustered 10 M similar payment terms across 9 Confidential businesses to identify the best payment terms for Confidential .
  • Used big data ETL tool Talend to connect to Confidential databases to acquire payment terms data. Used Python’s regular expressions to clean and parse the data, to generate Annual Percentage Index (API) value for each payment term and to identify the best payment term across different businesses for each vendor.
  • Stored clean data back in Confidential as well has Hadoop File System ( Confidential ) for archiving.
  • Compared 1.8M purchase orders with 50k Confidential ’s products to identify the potential savings for Confidential if all purchases are done internally from other Confidential companies.
  • Used Python and Fuzzywuzzy package for Levenshtein distance measure of comparing strings.

Confidential

Business Improvement Controller

Responsibilities:

  • Identified data trends, and used training, test & validation data for predictive models ( Confidential models).
  • Analyzed large data set related to passenger arrivals at 5-minute intervals to develop forecasting model using R using SARIMA models. Applied ARCH/GARCH models, stationarity test, white-noise test on this high-frequency data
  • Executed overall data aggregation/alignment & process improvement reporting for a business consisting of 10K people.
  • Led analytics projects at Emirates Airlines group airport operations division to drive productivity and service level improvement.
  • Used SAS enterprise guide for forecasting annual passenger growth through airport terminal for effective operations planning.
  • Developed framework for decision support systems for managing airport operations in real-time. This involved understanding the business processes, data sources, constraints and developing critical path method for identifying the critical flights and operations for effective management of operations.
  • Analyzed staffs’ leave award conditions across different divisions of airport operations and developed common rules for software application development for automated staff leave award scheduling system.
  • Designed dashboards for monthly review of Airport Operation’s organization performance.
  • Developed dashboard to identify the status of manpower in the Confidential . This dashboard will give a snapshot of current status of manpower, manpower at various stages of recruitment.
  • Lead member of the consulting team to conduct market analysis, review of terminal layouts for one of the airports in middle east
  • Developed forecasting model co-relating air cargo growth with GDP growth
  • Achieved +/-3% forecast accuracy of cargo volumes thereby improving air cargo terminals’ productivity
  • Achieved 67% reduction in cargo delivery time at one of the cargo terminals using six sigma techniques
  • Developed software requirement specifications for decision support systems (DSS) application for air cargo terminal operations monitoring and improving productivity
  • Successfully implemented Resource Management Systems for dynamic allocation of resources in baggage services division of Dubai airport.
  • Led and successfully implemented IT projects aimed at staff attendance automation projects

Confidential

Planning and Industrial Engineer

Responsibilities:

  • Applied support vector machine, random forest, logistic regression models using Python-scikit to classify wave forms and achieved an accuracy of 85%
  • Classified movie review texts as positive or negative using Naive-Bayes algorithm with a prediction accuracy of 83% using Python
  • Summarized 81000 Reuters news articles with unique keywords, by implementing Confidential in Python
  • Extracted features from lyrics, applied machine learning techniques to predict genre of a song, and achieved an accuracy of 87%
  • Hadoop course on Confidential As part of this course wrote pig and hive scripts to extract information from million song data set and Chicago employee data set.
  • Participated in hotel forecast demand competition on confidential and achieved MAPE of 0.25
  • Time Series Analysis of Passenger Numbers: Forecasted passenger numbers at 5-minute interval in R using SARIMA models. Applied ARCH/GARCH models, stationarity test, white-noise test on this high-frequency data
  • Data Acquisition - Job Aggregator: Developed job aggregator that grabs job details from various job sites and eliminates duplicates across multiple sites using beautifulsoup, lxml packages in Python

We'd love your feedback!