We provide IT Staff Augmentation Services!

Data Scientist Resume

3.00/5 (Submit Your Rating)

Bethesda, MD

PROFESSIONAL SUMMARY:

  • Data scientist with 6+ years of experience in transforming business requirements into actionable data models, prediction models and informative reporting solutions working in a variety of industries including Healthcare and Retail.
  • Expert in the Data Science process life cycle including Data Acquisition, Data Preparation, Data Manipulation, Feature Engineering, Machine Learning Algorithms, Validation, Visualization and Deployment.
  • Strong knowledge in Statistical methodologies such as Hypothesis Testing, Principal Component Analysis (PCA), Sampling Distributions, ANOVA, Chi - Square tests, Time Series, Factor Analysis, Discriminant Analysis.
  • Proficient in Python and its libraries such as Numpy, Pandas, Scikit-learn, Matplotlib and Seaborn.
  • Efficient in preprocessing data in Python using Visualization, Data cleaning, Correlation analysis, Imputations, Feature Selection, Scaling and Normalization, and Dimensionality Reduction methods.
  • Experienced in building various machine learning predictive models using algorithms such as Linear Regression, Logistic Regression, Naïve Bayes Classifier, Support Vector Machines (SVM), Neural Networks, KNN, K-means Clustering, Decision Trees, Ensemble methods (Random Forest, AdaBoost, Gradient Boosting and Bagging).
  • Knowledge in Text Mining, Topic Modelling, Association Rules, Sentiment Analysis, MarketBasket Analysis, Recommendation Systems, Natural Language Processing(NLP).
  • Knowledge on Time Series Analysis using AR, MA, ARIMA, GARCH and ARCH model.
  • Experienced in tuning models using Grid Search, Randomized Search, K-Fold Cross Validation.
  • Experience working with Big Data tools such as Hadoop - HDFS and MapReduce, HiveQL, Sqoop, Pig.
  • Extensive experience working with RDBMS such as SQL Server, MySQL, Oracle and NoSQL databases such as MongoDB, Cassandra, HBase.
  • Proficient in developing and designing ETL packages and reporting solutions using MS BI Suite (SSIS/SSRS).
  • Experience in building and publishing interactive reports and dashboards with design customizations based on the client requirements in Tableau, Looker, PowerBIand SSRS.
  • Proficient in data visualization tools such as Tableau, Python Matplotlib, Python Seaborn, R Shiny, R ggplot2 to create visually powerful and actionable interactive reports and dashboards.
  • Knowledge and experience working in Waterfall as well as Agile environments including the Scrum process and using Project Management tools like ProjectLibre, Jira/Confluence and version control tools such as Github/Git.

TECHNICAL SKILLS:

Databases: MS SQL Server, MongoDB 3.x, MySQL 5.x, Oracle, HBase

Statistical Methods: Hypothetical Testing, ANOVA, Time Series, Confidence Intervals, Bayes Law, Principal Component Analysis (PCA), Chi-square test, Chebyshev's inequality

Machine Learning: Linear Regressions, Logistic Regression, Na ve Bayes, Decision Trees, Random Forest, Support Vector Machine(SVM), Neural Networks, Sentiment Analysis, K-Means Clustering, K-nearest Neighbors (KNN), Ensemble Methods, Gradient Boosting Trees, Ada Boosting, PCA, LDA

Hadoop Ecosystem: Hadoop, MapReduce, Hive QL, HDFS, Pig

BI Reporting Tools: Tableau 10.x/ 9.x, MS SQL Server Integration Service and Reporting Service (SSIS/SSRS), Power BI

Data Visualization: Tableau, Python (MatPlotLib, Seaborn), R(ggplot2), Power BI, Qlikview

Languages: Python 2.x/3.x (Numpy, Pandas, Scikit-learn, Dask, Matplotlib, Seaborn), R (dplyr, ggplot2, rpart, caret, randomForest, gbm, trees, arules, h2o, neuralnet), SQL(MySQL), C, Matlab

Operating Systems: UNIX/UNIX Shell Scripting (via PuTTY client), Linux and Windows XP/7/8/10, Mac OS

Other tools and technologies: Azure ML Studio, Google TensorFlow, Apache Tomcat Webserver, MS Office Suite, Lucid Chart, StatTools, ProjectLibre, Google Analytics, Google Tag Manager, Salesforce, MS SharePoint, Trello, JIRA, Confluence, Github/Git, AWS (EC2/S3/Redshift/EMR/Lambda)

PROFESSIONAL EXPERIENCE:

Confidential, Bethesda, MD

Data Scientist

Responsibilities:

  • Worked closely with University leadership and staff to understand their analytical needs and develop/maintain models that provide insight into student success and operational efficiency.
  • Translated project needs and goals into a plan describing research questions, data sources, endpoint definitions, and the advanced analytics approach, using rigorous statistical methods and machine-learning techniques
  • Used data-driven and statistically valid methodologies to develop and maintain models that in corporate key internal and external metrics to produce short and long-range models.
  • Extensively involved in all phases of data acquisition, data collection, data cleaning, model development, model validation, and visualization to deliver data science solutions.
  • Worked on data cleaning and ensured data quality, consistency, integrity using Pandas, Numpy.
  • Tackled highly imbalanced datasets using sampling techniques like undersampling and oversampling with SMOTE (Synthetic Minority Over-Sampling Technique) using Python Scikit-learn.
  • Used PCA and other feature engineering techniques to reduce the high dimensional data, feature normalization techniques and label encoding with Scikit-learn library in Python.
  • Used Pandas, Numpy, Seaborn, Matplotlib, Scikit-learn in Python for developing various machine learning models such as Logistic regression, KNN and Gradient Boosting.
  • Worked on Amazon Web Services cloud services to do machine learning on big data.
  • Used cross-validation to test the models with different batches of data to optimize the models and prevent overfitting.
  • Experimented with Ensemble methods to increase the accuracy of the training model with different Bagging and Boosting methods.
  • Deployed the model on AWS EC2 using Flask.
  • Collaborated with other analysts to identify underlying trends (both internal and external) that impact current and future institutional goals.
  • Created data visualizations to provide insight to complex business problems.
  • Supported with the development and testing of machine learning algorithms along with creating reports to display the status and performance of deployed model and algorithms with Tableau.

Confidential, Dallas, TX

Data Scientist/ Modeler

Responsibilities:

  • Derived key insights from big data and delivered strategic insights for improved decisions.
  • Developed personalized offers generation by analyzing and merging data from transactional and hierarchical datasets that includes purchases through instore, website and mobile application.
  • Worked with the Business Intelligence team to implement data strategies, build data flows and develop conceptual data models.
  • Imported data from various data sources, performed transformations using Hive, and loaded data into HDFS.
  • Created logical and physical data models using best practices to ensure high data quality and reduced redundancy specifically in a Qlik environment.
  • Optimized and updated logical and physical data models to support new and existing projects.
  • Maintained conceptual, logical and physical data models along with corresponding metadata.
  • Developed data models according to company standards and best coding practices to ensure consistency of data models.
  • Developed different Machine algorithms such as Logistic Regression, SVM, Decision trees, Random Forests, XGBoost to predict customer insight, target marketing, potential lapse customers.
  • Predicted the probability for retention of all the existing member customers using Logistic regression and XGBoost models.
  • Developed a model to forecast the demand that optimized the understocking-overstocking problem by analyzing the stock & sales data at a weekly level.
  • Performed visualizations using Tableau, Matplotlib, Seaborn and Plotly.
  • Recommended opportunities for reuse of data models in new environments.
  • Performed reverse engineering of physical data models from databases and SQL scripts.
  • Evaluated data models and physical databases for variances and discrepancies.
  • Validated business data objects for accuracy and completeness

Confidential

Data Analyst

Responsibilities:

  • Used Python to generate graphs for given data set features.
  • Experience working with probability distributions in Python.
  • Machine learning for regression, Classification, NLP, information retrieval and anomaly detection.
  • Created end-end machine learning pipeline.
  • Used Python to find customers who are interested in buying the products and who need customer service.
  • Practical Knowledge of machine learning and software development and a thorough understanding of the underlying mathematical concepts.
  • Strong quantitative analysis and data visualization skills.
  • Statistical Skills: ANOVA, regressions, generalized linear models, time series.
  • Extensive knowledge in data mining and predictive modelling: linear and logistic regression, decision trees, random forest, K- nearest neighbors, SVM, clustering.
  • Big data analysis and predictive modelling using MapReduce and HIVE.
  • Extracted business insights from massive information and communicate effectively.
  • Self-driven, Passionate about creating data driven solutions that fuel business growth.

Confidential

Data analyst

Responsibilities:

  • Developed ETL processes for data conversions and construction of data warehouse using IBM InfoSphere DataStage.
  • Used Star Schema and designed Mappings between sources to operational staging targets.
  • Involved in defining the business/transformation rules applied for sales and service data.
  • Define the list codes and code conversions between the source systems and the data mart.
  • Provided On-call Support for the project and gave a knowledge transfer for the clients.
  • Used Rational Application Developer (RAD) for version control.
  • Developed transformations using jobs like Filter, Join, Lookup, Merge, Hashed file, Aggregator, Transformer and Dataset.
  • Worked with internal architects and, assisting in the development of current and target state data architectures.
  • Coordinate with the business users in providing appropriate, effective and efficient way to design the new reporting needs based on the user with the existing functionality.
  • Remain knowledgeable in all areas of business operations to identify systems needs and requirements.
  • Document the complete process flow to describe program development, logic, testing, and implementation, application integration, coding.

We'd love your feedback!