We provide IT Staff Augmentation Services!

Data Scientist Resume

5.00/5 (Submit Your Rating)

Tampa, FloridA

SUMMARY:

  • Hands on SparkMlib utilities such as including classification, regression, clustering, collaborative filtering, dimensionality reduction.
  • Experience in implementing NaiveBayes and skilled in Random Forests , Decision Trees , Linear and Logistic Regression , SVM, Clustering , neural networks, Principle Component Analysis and good knowledge on Recommender Systems.
  • Worked on implementation of Dynamic programming in Machine learning as Reinforcement learning.
  • Worked on different libraries related to Data science and Machine learning like Scikit - learn, Opencv, numpy, Scipy, Matplotlib, pandas, Json, SQL etc.
  • Also worked on different python web frameworks like Django.
  • Proficient in statistical programming languages like R and Python including Big Data technologies like Hadoop, Hive .
  • Involved in all the phases of project life cycle including Data acquisition (sampling methods: SRS/stratified/cluster/systematic/multistage), Power Analysis, Hypothesis testing, EDA (Univariate & Multivariate analysis), Data cleaning, Data Imputation (outlier detection via chi square detection, residual analysis, PCA analysis, multivariate outlier detection), Data Transformation, Features scaling, Features engineering, Statistical modeling both linear and nonlinear (logistic, linear, Naïve Bayes, decision trees, Random forest, neural networks, SVM, clustering, KNN), Dimensionality reduction using Principal Component Analysis (PCA) and Factor Analysis, testing and validation using ROC plot, K- fold cross validation, statistical significance testing, Data visualization.
  • Having Strong years of experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
  • Also have an ability to design a Data warehouse/Mart or a Database model on different platforms like SQL, NoSQL, Oracle, SQL Server with full redundancy and normalization.

TECHNICAL SKILLS:

  • Python with OpenCV, Django web framework, Scikit-learn, pandas, json, SQL, TSQL and Big Data Platforms like Spark and Hadoop.
  • Java
  • PHP

PROFESSIONAL EXPERIENCE:

Confidential, Tampa, Florida

Data Scientist

Responsibilites:

  • The present project is to build an algorithm that accurately classifies credit card holders among multiple classes based on the historical data available on multiple variables. Further, the aim was to improve bank's efficiency by reducing default rate and speed up the process while offering new products.
  • Also involved in a project to identify the employees' access level, based on their current & historical tasks and duties.
  • Developed data solutions to support strategic initiatives, improve internal processes, and assist with strategic decision-making and design SWOT Analysis.
  • Extraction by developing a pipeline using Hive (HQL) to retrieve the data from Hadoop cluster, SQL to retrieve data from Oracle database and used ETL for data transformation .
  • Performed Data Cleaning, features scaling, features engineering using pandas and numpy packages in python.
  • Replacement of missing data and perform a proper EDA, Univariate and bi-variate analysis to understand the intrinsic effect/combined effects.
  • Involved working on different databases like json, SPARK/HADOOP, XML, NoSQL and SQL of different platforms etc.
  • Involved working in Data science on different data transformation and validation techniques like Dimensionality reduction using Principal Component Analysis (PCA) and Factor Analysis, testing and validation using ROC plot, K- fold cross validation, statistical significance testing.
  • Evaluated models for feature selection and elastic technologies like Elasticsearch, Kibana etc.
  • Used predictive analytics and machine learning algorithms to forecast key metrics in the form of designed dashboards on to AWS and Django platform for the company’s core business.
  • Provided data and analytical support for the company’s highest-priority initiatives.
  • Used Python / R to design many other machine learning algorithms such as Decision Tree, linear regression, multivariate regression , Naive Bayes , Random Forests , K-means , & KNN based on Unsupervised/Supervised Model that help in decision making.
  • Generated visualizations using Tableau and R-Shiny to present the findings on call center analytics.
  • Implementation of Reinforcement learning techniques in the field of Machine learning by following Dynamic programming using Python.
  • Also worked with several R packages including knitr, dplyr, SparkR, CausalInfer, spacetime .

Environment: Python 2.x/3.x, R, CDH5, HDFS, Hadoop 2.3, Hive, Linux, Spark, Tableau Desktop, SQL Server 2012, Microsoft Excel, Matlab, Spark SQL, Pyspark, SQL, Scikit-learn, Pandas, AWS, XML, json.

Confidential, North Charleston, South Carolina

Data Scientist

Responsibilities:

  • Involved in an ecommerce project to analyze the insights of customers. Perform Sentiment Analysis on the customers reviews about different products, study the customer satisfaction based on the purchase made.
  • Also analyze the category and product based sales based on the customer demographics and support the product promotion team to increase the sale.
  • Understanding and implementing the process of MapReduce using various Big Data platforms like Hadoop/Spark SQL API in python.
  • Etraction of data from different database and warehouses and transforming them as per the requirements.
  • Performed data profiling to learn about the behavior of various features and finding dependencies with in them.
  • Understanding the complexity in available raw dataset and came up with different solutions like Dimensional reducing and Feature scaling to transform the dataset that fit to a best analytical model depending on the perspective of study.
  • Performed data cleaning and feature selection using MLlib package in PySpark and working with deep learning frameworks such as Caffe, Neon etc.
  • Design of Dashboards showing Categories and product based reports with key performance indicators.
  • Application of various machine learning algorithms and statistical modeling like decision trees , text analytics, natural language processing (NLP), supervised and unsupervised , regression models , social network analysis, neural networks, deep learning , SVM , clustering to identify Volume using scikit-learn package in python , Matlab.
  • Predict the impact of various promotional decision makings on overall sale of the Company using various Clustering techniques of Machine learning approaches like K-means, Maximum Likelihood and other Hierarchical Clustering techniques.
  • Maintain the robustness and effective fit to make predictions.
  • Responding and coordinating to making changes to the existing Data Model based on ad hoc analytical reports

Environment: Python/R, Bigdata Hadoop/Spark SQL, Scikit-learn, Pandas

Confidential, Detroit, Michigan

Data Analyst

Responsibilities:

  • Worked on Housing loans, in finding the Prospects, analyze their purchase power that helps the Sales team to convert the prospect into a sale and predict the loan defaulters by risk assessment estimations.
  • Quires on different databases related to structural and non-structural data and worked on its extraction.
  • Parsing data, producing concise conclusions from raw data in a clean, well-structured and easily maintainable format using Python.
  • Making of power analysis to determine the sample strength in the datasets.
  • Extracting the hidden information from different quantitative variable by variable analysis and analytics.
  • Analyzing the data set and explore different dimensions/patterns in it, for Feature Engineering.
  • Involved with a team of developers on core predictive models like linear regression , logistic regression and KNN etc. identifies the likelihoods, to improve the predictive performances and predict the binaries.
  • Creating various B2B Predictive and descriptive analytics on to the dashboards.
  • Worked with Business intelligence tools and visualization tools such as Tableau, ChartIO, etc.
  • Deployed GUI pages using JSP/PHP, HTML, DHTML, XHTML, CSS, JavaScript and Ajax.

Environment: Python, SQL, Machine Learning, XML, JSP, PHP, Pandas, R

Confidential

Data Analyst

Resposibilites:

  • Study and analyze the months on months financial reports and generate an analysis report, finding the interesting facts that helps the business model using Matlab.
  • Monitored the costs and trend analysis by providing variances to respective stakeholders.
  • Prepared annual budget and periodic forecasts.
  • Understand the requirements and provide the reports as per the requirements.
  • Design a solution for the areas of concern by exploring the facts from the available data.
  • Provide ad-hoc reports as per the requirements by conducting peer reviews and meeting with the client.
  • Conduct feasibility study, impact analysis and study of case scenarios.
  • Quired the database with complex quires of SQL Server, MySQL and Oracle platforms.

We'd love your feedback!