We provide IT Staff Augmentation Services!

Data Scientist Resume

3.00/5 (Submit Your Rating)

Seattle, WA

SUMMARY:

  • Over 5+ years of Professional experience in Machine Learning, Deep Learning, Natural Language Processing(NLP), Artificial Intelligence(AI), Data Mining, Data Analysis.
  • Hands on experience in implementing Machine Learning algorithms like Linear Regression, K - Nearest Neighbors(KNN), Logistic Regression, Naïve Bayes, Decision Trees, Random Forests, SVM (Support Vector Machines), Gradient Boosted Decision Trees using XGBOOST.
  • Performed Exploratory data analysis (EDA) extensively to know the data insights and also to know the important features in the data, worked with lot of visualization tools like MatPlotLib, seaborn, ggplot, pygal, experienced in implementing machine learning programs in Python, R, Scala.
  • Expertise in unsupervised learning algorithms like K-Means, Density Based clustering(DBSCAN) and Hierarchical clustering and decent experience on Recommender systems.
  • Experienced in Dimensionality reduction techniques such as Principal Component Analysis(PCA), Confidential - distributed stochastic neighbor embedding ( Confidential -SNE) and Truncated SVD.
  • Expert level Mathematical and statistical knowledge on Geometry, Statistics Probability, Linear Algebra, Integration, Differentiation and Proficient in understanding distributions and their corresponding properties.
  • Proficient in deep learning techniques like Deep Neural networks, Artificial neural networks, implemented neural computation, Convolutional neural nets, knowledge of using GPU in parallel computing approaches.
  • Decent Experience in building Deep learning models with frameworks like Tensor Flow and keras, used SAS for developing algorithms.
  • Hands on experience in applying Optimization techniques like Gradient Descent and Stochastic Gradient Descent.
  • Experience in working with R, SAS and also in Hadoop environment. Good command on Linux environment, experience in working with customer data.
  • Worked and extracted data from various database sources like Oracle, SQL Server, DB2, and Teradata and experienced in writing complex SQL queries like stored procedures, joins, triggers and subqueries.
  • Experience with data visualization using Tableau to publish and present dashboards, storyline on both web and desktop platforms.
  • Good hands-on experience and proficient with structured, unstructured and semi-structured data using a broad range of data science programming languages and big data tools including Python, SQL, Spark ML lib and Scikit-Learn, Experience in using GIT Version Control System. Accessed data via variety of API / RESTful services.
  • Familiar with Google and other cloud Machine Learning, experience with data management.
  • Excellent communication skills and a Conversationalist, I speak the language of statistics and I can explain what I’m doing to my grandma (non-technical guys).
  • An Innovator, I will always keep asking “Why”, strong passion towards Machine learning(ML) and Artificial Intelligence(AI).

PROFESSIONAL EXPERIENCE:

Confidential, Seattle, WA

Data scientist

Responsibilities:

  • Full stack experience in SDLC which involves data collection, analyzing data, Visualization, automation.
  • Converting the data to desired formats that helps us in data cleaning and data preprocessing.
  • Performed EDA (Exploratory data analysis) to know the insights of the data.
  • Extensively worked on cleaning, preprocessing the data and dimensionality reduction using techniques like PCA (principal Component Analysis) and Confidential -SNE ( Confidential - stochastic neighborhood embedding) to reduce the dimensions of the higher-dimensional data.
  • Performed feature engineering, performed NLP by using some techniques like Word2Vec, BOW (Bag of Words), tf-idf, Avg-Word2Vec, if-idf Weighted Word2Vec.
  • Performed data mining on un-structured e-mail text data, by NLP (Natural Language Processing) and some of the deep learning techniques are also used.
  • Used Naïve bays classifier to train the reviews dataset.
  • Built frequency distribution for all words and frequency distribution for words within positive and negative labels.
  • Developed Web Services for online text polarity classification.
  • Deploying and maintaining sentiment analysis web app on AWS.
  • Visualized, interpret, found the reports and develop uses of data by python Libraries like Pandas, Numpy, Scikit-learn, MatPlotLib, Seaborn.

Environment: Python, SQL, Oracle 12c, NLTK, Recurrent Neural Networks, LSTM cells, Natural Language Toolkit, NumPy, SciPy, Pandas, Matplotlib, Seaborn, Scikit-Learn, Tensor Flow, Keras.

Confidential, Richmond, VA

Data Scientist

Responsibilities:

  • Involved in data cleaning and pre-processing of this highly imbalanced dataset. Splitting the dataset into training and test set. Trained a random-forest model and validated for the identification of fraud activity. This include challenges consisting of 205 frauds in a million transactions, and cap false positive rate to 2% owing to manpower limitations.
  • Active participation in researching and approaching SME’s (Subject matter experts) for the identification of fraud indications. Developed features based on online activity, Card transactions, updating personal data, travel transactions, past activity, etc.
  • Worked with imbalanced dataset using some of the sampling techniques like up-sampling, down-sampling and SMOTE (Synthetic minority over- sampling technique) using Scikit-learn.
  • Extensively used dimensionality reduction techniques like PCA and Confidential -SNE as a part of feature engineering. Applied various machine learning algorithms like regression models, clustering, and Support vector machines(SVM) using Sickit-learn in python.
  • Able to predict with an accuracy of 95% with a false positive rate of 2.5%, comparatively larger with the existing statistics.
  • Implemented a Python-based random forest distribution via PySpark and MLlib.
  • The performance is measured using log-loss function used AUC and ROC curves for the feature selection.
  • Created and maintained reports for visualization of the status and performance of the deployed model with Tableau.

Environment: NumPy, Pandas, Matplotlib, Seaborn, Scikit-Learn, Tableau, SQL, Linux, Git, Microsoft Excel, Spark SQL, Logistic Regression, Random Forests, Decision Trees, Confidential -SNE, PCA, Tensor Flow, K-Means, Natural Language Tool Kit.

Confidential, St Louis, MO

Data Analyst

Responsibilities:

  • Expert in data validation, cleansing, consolidation and mining for sense, consistency and accuracy and compliance standards.
  • Captured data lineage for all the top-level reports by validating the authorized data sources with system of records and system of origin.
  • Performed analysis on existing data model to understand the methodology of model and help QlikView developers understand the requirements.
  • Created scheduled jobs for QVD extracts and report reloads using Qlikview Server and publisher.
  • Validated the data at QVD level and Dashboards level using Qlikview.
  • Provided tech team support for SDLC Data Warehouse and assisted Project Managers in establishing plans, risk assessments and milestone deliverables.
  • Designed data models using Oracle Designer and designed programs for data extraction and loading into Oracle database and managed database tables, procedures and indexes.
  • Created SQL-Loader scripts to load legacy data into Oracle staging tables and wrote SQL queries to perform Data Validation and Data Integrity testing.
  • Implemented enhancements to existing software products.

Environment: Pivot tables, UNIX, QVD, MS Excel, Teradata, AWS, Python, SAS, SQL, Oracle.

Confidential

Data Analyst

Responsibilities:

  • Database management, maintenance and data analysis, processing and testing.
  • Tested and ensured data accuracy through the creation and implementation of data integrity queries, debugging and trouble shooting.
  • Worked with business analysts for understanding the problem statement and their requirements.
  • Extracted data from various relational databases and performed SQL queries depending on how the data needs to be modified, Used FTP to download SAS formatted data.
  • Developing new code and modifying existing code to extract data from various data sources like DB2, oracle, Used MS-excel and SQL server extensively to manipulate the data for business requirements.
  • Worked on creating new datasets from uncleaned or raw data using some importing techniques and also modifies the datasets which already existed using join, set, sort, merge, update and some other conditional statements.
  • Closely worked with Machine learning engineers to analyze the data based upon their requirements.
  • Experienced in creating pivot tables for analyzing data in excel.
  • Hands-on experience in data analyzing and wrote MySQL queries for improved performance.
  • Optimized queries with some manipulations and modifications in MySQL code and removed unwanted columns and duplicate data.
  • Exceeded expectations by helping machine learning team in data cleaning and preprocessing which is used for building their machine learning algorithms.
  • Decent experience in creating dashboards in tableau for report submissions.

Environment: Base SAS v9.2, SAS/Macros, Microsoft Excel, SAS/SQL, SAS Enterprise Guide v4.2, SAS/ACCESS, Mainframes, UNIX, Teradata.

We'd love your feedback!