We provide IT Staff Augmentation Services!

Data Scientist Consultant Resume

3.00/5 (Submit Your Rating)

Waltham, MA

SUMMARY

  • Overall 8 years’ experience as a Data Scientist/Analytics in delivering end - to-end advanced analytics and BI solutions using Statistics, Mathematical Modeling, and Machine Learning. Strong Project Management skills and experience leading/collaborating with cross-functional teams.
  • Experience in transforming raw data into actionable strategic knowledge to gain insight into business processes, and thereby guide and influence businesses in their decision-making.
  • Solid experience processing massive amounts of structured and unstructured data using Spark/SQL/Hive.
  • Proficiency in Python for numerical/statistical programming (including Numpy, Pandas, and Scikit-learn)
  • Proficient in building NLP pipeline using Apache Spark for Massively Parallel NLP use cases.
  • Excellent knowledge and experience with Hadoop Architecture and other components of its ecosystems like HDFS, YARN, Map Reduce, Hive, Spark.
  • Proficiency inSpark using Scala for loading data from the local file systems like HDFS, Amazon S3, Relational and NoSQL databases usingSpark SQL and Import data into RDD.
  • Experienced in performing in memory data processing for batch, real time, and advanced analytic using Apache Spark (Spark SQL &Spark-Shell).
  • Solid understanding of advanced supervised and unsupervised machine learning models (e.g., auto-encoders, convolutional networks, XGBoost, etc.).
  • Understanding of training multi-layer Neural Network for deep learning using Spark and TensorFlow.
  • Experience developing dynamic BI dashboards using Tableau and presenting to the stakeholders.
  • Dexterous in creating compelling data visualization using ggplot2/Plotly/D3.js/DECK.GL.
  • Experience in solving Optimization problems using Apache OpenOffice and Julia programming.
  • Extensive experience in in-depth data analysis on different databases and structures. Strong knowledge in writing SQL Queries, sub-queries, and joins.
  • Skilled in providing analytic support including data importing, data wrangling and data visualization.
  • Significant experience with CRM databases/systems, to query and analyze customer data. Plus, experience with Google Analytics, Tableau, MS Excel and SQL Server Reporting Services (SSRS).
  • Proficient in database development in RDBMs: Oracle, PostgreSQL, MySQL, and MS-SQL.
  • Excellent knowledge in creating Databases, Tables, Stored Procedure, DDL/DML Triggers, Views, User defined data types,effective functions,Cursors and Indexes.
  • Extensive knowledge of financial accounting, corporate finance, investment banking, fixed income products, Dodd-Frank act and American/European financial systems.
  • Experience and profound understanding in digital marketing: SEO, SEM, PPC, Adwords.
  • Experienced in developing test plans, cases, scenarios and strategies as per the business requirements to match the functional requirements and UML diagrams.
  • Solid experience in Software Development Life Cycle (SDLC) including requirement gathering, analysis, design, development and testing in both Waterfall and Agile methodologies.
  • Extraordinary presentation skills and ability to communicate with both tech/non tech stakeholders.
  • Exceptional problem-solver and fast learner with insatiable curiosity to learn new technologies.
  • Expertise in using APIs.

TECHNICAL SKILLS

Project Management: Agile Scrum, JIRA, SMARTSHEET, MS-Project

Programming: R, Python, Java, C++ (Familiarity: Scala, Julia)

Tools: PyCharm, Eclipse, IPython Notebook/Jupyter, Apache Zeppelin, Git, Adobe Analytics

MapReduce/Data Pipeline: Hadoop/HDFS, Java MR, Python Streaming MR, Hive-Python plugins, Spark, Splunk

Cloud: Hadoop-on-Azure, AWS/EMR/EC2/S3 (also direct-Hadop-EC2 (nonEMR))

SQL/NoSQL: Oracle, MySQL, PostgreSQL, Hive, Hydra, Spark SQL (Familiarity: Cassandra, Presto)

BI Tools: Tableau 10.0, PARSE.LY, Ploty, Oracle BI, Kibana, CLOUDVIEW

APIs: LinkedIn, Twitter, Open Layers, RESTful API

ML/DL Toolkit: Scikit-learn, Spark MLib, Intel BigDL, Keras, TensorFlow, Theano, Torch, H2O, CUDA

ML Algorithms: Supervised: Multivariate Regression, Logistic Regression, Time-Series, Decision Trees, Naïve Bayes Random Forest, XGboost, k-Nearest Neighbors, Support Vector Machine Unsupervised: Clustering (K-means, Hierarchical), Recommendation Systems (Content/Collaborative Filtering and Hybrid Models), Neural Networks, Deep Learning

Web: HTML CSS, PHP, JQuery, JavaScript, D3.js, Three.js

PROFESSIONAL EXPERIENCE

Confidential, Waltham MA

Data Scientist Consultant

Responsibilities:

  • Extracted and transformed customers and marketing datasets obtained from multiple sources (Oracle, PostgreSQL), loading them into EXALEAD for indexing, and developed a web based front end/search bases applications using a combination of Grails/Java + JQuery + d3.js.
  • Developed a few dynamic BI dashboard and produced regular reporting to track key KPIs, sales and performance matrixes across multiple channels, used Tableau and D3.js.
  • Analyzed a vast amount of text data from various sources for NLP use cases, used Apache Spark for Massively Parallel NLP and Logistic Regression to predict the probability of the survey questions.
  • Extracted data from RDBMS database to Hive using Squoop and performed advanced Hive SQL.
  • Implemented Spark scripts using Spark and Spark SQL to access Hive tables into spark for faster processing of data.
  • Built the response model for email marketing by using XGBoost-Spark - 350% increase response rate.
  • Built and tested Time Series model in Spark R for trends and predictive sales pipeline analytics, used various smoothing methods to improve the model, and integrated the algorithm with the sales dashboard using D3.js, resulting 98.5% accuracy in FY17Q1.
  • Optimized digital marketing budget by analyzing AdWords data using Apache OpenOffice, saving 17% on budget with 750% ROI on PPC channel.
  • Generated hypothesis for A/B testing for Banner Ads on google, additionally, for multiple alternatives, choose multi-armed bandit algorithm to create more value in spending.
Environment: Hadoop Hive Spark, Python, R, Java, SQL, EXALEAD Cloudview, SQL, HTML, D3.js, JQuery

Confidential, Waltham MA

Data Mining Scientist

Responsibilities:

  • Created a product portfolio application using previous years of revenue data, EXALEAD for indexing and HTML CSS JavaScript and Java to create an application.
  • Finding outliers, errors, trends, missing values, and distribution in the data. Utilized techniques like Histogram, Bar plot, pie-chart, scatter chart, box plots to determine the initial conditions of the data.
  • Loaded relational databases to Apache Spark using Python, then performed Time-Series analysis.
  • Built a multivariate regression model for revenue and channel analytics with detailed correlation and causation analysis, tested and compared with variable selection regression method such as LASSO, Ridge and Elastic Net using Spark MLib, integrated with global marketing dashboard using JS.
  • Segmented customers for advocacy marketing using K-mean model, built and tested the model in Spark Mlib and used D3.js for visualization.
  • Analyzed usage trends for content with Adobe analytics and CRM Data Extracts with Siebel Analytics.
  • Analyzed customer’s feedback data using text mining package spark-ts in Spark, helping the team to know the sentiments of the customers about the product, ENOVIA.
  • Improved UX|UI of ENOVIA’s webpage by collaborating with the design and engineering team.
  • Developed and maintained the BI dashboard and produced regular reporting to track key KPIs, sales and performance matrixes across multiple channels, used Tableau and D3.js.
  • Developed various data visualization dashboards comprising campaign spending, performance matrix and ROI using Tableau and Shiny.

Environment: SparkMLib, SQL (Hive, Spark), Python, JavaScript, Tableau, R Shiny, EXALEAD, APIs

Confidential, Needham MA

Data Analytics/BI Scientist, Marketing

Responsibilities:

  • Built a real-time streaming model using Spark Streaming - Loaded streaming data from Amazon S3 to Spark RDDs, wrote a program to take the live streaming of Tweets, built a machine learning model using Spark MLlib TF-IDF in Scala, invoked it from SQL to build a dashboard that uses a live uses this machine learning model and filter out things for users.
  • Spatial data analysis and geo clustering using spatial libraries (sp, rgdal, maptools, spatstat, dbscan) - combining geo location data and search history to build recommendation systems.
  • Built classification models, typically SVM, using ksvm in e1071 package for a variety of use cases.
  • Determined the missing data, outlier and invalid data and applied appropriate data management techniques.
  • Worked on data manipulation and raw marketing data of different formats from multiple sources and prepared the data for further analysis using Reshape2 and ggplot2 packages.
  • Analyzed different trends on historical data and developed a web-app using Shiny present the findings.
  • Designed custom reports, charts, tables and dashboards using Tableau/Shiny to help the marketing and operations teams in their decision making process.
  • Create focused reports and dashboards on content/channel performance, lead generation and conversion rates.
  • Strategic Consulting, including business plan and marketing strategy development.
  • Wrote simple and advanced SQL queries and scripts to create standard reports for senior managers.

Environment: R Studio, R Shiny Server, Hadoop, Hive, MySQL, CRM, Tableau, R Markdown, HTML, CSS, JS.

Confidential

Data Scientist/Project Manager

Responsibilities:

  • Worked with different teams (survey, pharmaceutical and research and development, claims etc.) to gain insights about the data concepts behind their health symptoms and business.
  • Analyze customer, behavior data, symptoms data, transaction data and campaign data to identify trends and patterns of data in different visualization techniques like Seaborn library in Python.
  • Explored connections between diseases and drugs in an unsupervised way using Python, Spark MLib.
  • Extraction of large amounts ofdata for analysis and reporting. Responsible for documentation of all analysis as well asdata discrepancies.
  • Working closes with other analysts to reconcile all issues related todata production, extraction and delivery in order to ensure the integrity of thedata and the reporting that it is used for.
  • Involved in initial data pattern recognition and data cleaning using Numpy and Pandas in Python.
  • Determined the missing data, outlier and invalid data and applied appropriate data management techniques.
  • Worked on data manipulation and raw marketing data of different formats from multiple sources and prepared the data for Sentiment analysis of all the customer medical issue data using packages like NLTK (Natural Language Processing with Python / Analyzing Text with the Natural Language Toolkit).
  • Wrote script in python to predict number of people getting effect of some diseases, by collecting set of predicted (symptoms) data from all medical sectors and evaluated with outcome data and Make the aware of people.
  • Wrote simple and advanced SQL queries and scripts to create standard reports for senior managers.
  • Designed custom reports, charts, tables and dashboards using Tableau and for marketing and operations teams for their decision-making process.
  • Implement and customize web analytics and website optimization tools based on business needs, specifically Google Analytics.
  • Analyzed generic search data from GoogleAdwords to see the top trend pertaining to child health, getting the idea what the general population is looking for.
  • Analyzed patient’s data in R from BCH Emergency Room Data to identify the most common illnesses (complaints) in children and correlation between complaints and billing data, and used ggplot to visualize charts and trends, findings were used in internal research paper which paved the way for managers to make sound decisions and create project strategy.
  • Conceptualized and normalized a database for clinical researchers to access the appropriate health care information.
  • Used automated testing for applications and validation methods for analytical models.
  • Collaborated with Amazon Alexa team to partner with KidsMD platform.

Environment: Spyder studio, PYTHON, Sentiment analysis, Time series Analysis, Google Analytics, My SQL, CRM, Tableau, MS Access.

Confidential, Boston, MA

Data Analyst

Responsibilities:

  • Analyzed business requirements, system requirements, data mapping requirement specifications, and responsible for documenting functional requirements and supplementary requirements.
  • Initial pattern recognition and data cleansing using dpylr package in R.
  • Worked on data manipulation on raw purchasing data in different formats from multiple sources (oracle, byways) and prepared the data for further analysis using ggvis and ggplot2 packages.
  • Web scrapping using Python and API’s for commodity’s price comparison.
  • Used multivariate regression to predict spending at the departmental and categorical level.
  • Prepared annual/semi-annual/quarterly/monthly reports for the department using SQL server reporting service (SSRS) and advanced excel functions like Vlookups, PivotTables, Merging, Sorting.
  • Analyzed reports of data duplicates or other errors based on monthly or daily data reports.
  • Provided visual spend analytics using Tableau to identify opportunities to reduce cost, track contract compliances, and measure supplier’s performance - 11% cost reduction.
  • Strategic contributions to finance and contract and compliance board meetings regarding optimization of university’s budget and new projects.
  • Update and maintain the custodian database in PeopleSoft Finance for asset management.
  • Analyze reports of data duplicates or other errors based on monthly or daily data reports.
  • Conduct workshops to train new hires about PeopleSoft Finance tool.

Environment: Excel (Vlookups, PivotTables etc), Tableau, PeopleSoft, SQL Server, R Shiny, R Studio, Python

Confidential

Research Marketing Analyst

Responsibilities:

  • Executed pay-per-click (PPC) advertising plan by analyzing AdWords data extracted form Keyword Planner and third party sources.
  • Analyzed data from the third party using R to create a marketing plan Snapchat platform.
  • Optimized budget for digital marketing using linear programming in OpenSolver resulting 20% reduction in budget from previous year.
  • Developed and executed SEO strategy that achieved and sustained top 5 ranking on Google and Yahoo

Environment: R, GoogleAdwords, Apache OpenSolver, Google Analytics.

We'd love your feedback!