We provide IT Staff Augmentation Services!

Data Scientist Resume

4.00/5 (Submit Your Rating)

Dallas, TexaS

SUMMARY

  • Principal data Scientist with 12 years of extensive experience in developing & delivering end - to-end data science, Machine learnings products and solutions.
  • Identified DS/ML use cases that creates immediate benefits to business. Lead highly innovative and vibrant team of Data Scientists, Data engineers in developing the next generation AI/ML enabled solutions.
  • Build image classifier, NLP solutions, forecasting model, Recommender systems, Personalization solutions, Strong knowledge in Deep Learning, Statistical Analysis, Machine Learning, Hypothesis testing & A/B testing.
  • Have analyzed & reported actionable insights & trends from Big Data and presented to the executives, strategic team, project sponsors in business-friendly terms. Have closely worked with Business & IT teams to create analytics workbench.

TECHNICAL SKILLS

Machine Learning & Statistics: Recommender System (Deep & Wide Neural Networks, Collaborative Filtering, Matrix Factorization),Time series forecasting, Document clustering, Latent Semantic Analysis, Model Evaluation (MAE, MAPE, MASE), CNN, RNN-LSTM, Attention Networks, NLP,NLU, Neural Networks, Boosted Trees, Regression, Convolutional Net, Word2Vec, Gradient Boosting, Classification, Regression, Random Forest, Clustering(K-means), SVM, Social Media Analytics, Sentimental analysis, Bagging, Boosting, Word Embedding, Text summarization, Text parsing, TF-IDF, End to end model building, develop and deploy Predictive model building (Developed more than 50 models in various projects),Voice Biometrics, Intelligent Process Automation, Text labelling and Categorization

ML & Data Science Libraries: Tensor Flow, Keras, Caffe, Pandas, Data.table, Forecast Hybrid OpenCV (image processing), Ggplot2, Caret, TM, Purrrr, NLTK (Natural Language Processing), Mahout, NER, AWS

Machine Learning Models: CNN, RESNET, Memory Network, RNN LSTM, RNN GRU, CONV1D, Random Forest, Gradient Boosting, GLM, ALS Matrix Factorization, Collaborative Filtering, LSA, PCA, SVD

Languages: R, Python, SQL, Linux

Tools: Tableau, R Studio, Jupytr Notebook, HDFS Hadoop, SAP HANA, SPARK, Microsoft Azure, MS Visual Studio, TFS, SharePoint, Clear Quest, HSD, Rally, Microsoft office, H20, GCP Certification(currently enrolled or in process)

PROFESSIONAL EXPERIENCE

Data Scientist

Confidential, Dallas, Texas

Responsibilities:

  • Worked on, multiple projects to leverage statistical learning/ML algorithms to automate Alternate Asset Servicing. The automation helped Arvos and Molvis to reduce errors and improve operational efficiency.
  • Further, developed BI reports that provided predictive analytics & Reporting with dashboards on analyst performance, client activity, future workload & anomalies based on data collected from discrete applications. Apache spark is used to develop ML models, API keys and various applications during projects.
  • Defined Project Scope, project Charter & Business Case Prototype ML algorithm for POC (Proof Of Concept) Performed Data Cleaning, features scaling, features engineering, Developed predictive models for use in machine learning platform using the scikit-learn python framework Improved statistical models using learning curves, parameter curves, feature selection, and regularization.
  • Performed Principal Component Analysis (PCA) & Linear Discriminate Analysis(LDA) Fine-tuned low bias & High variance trade off Defined the technical requirements of the analytic solutions.
  • Defined the data requirements of the analytic solution. Lead a team of Data Scientists/ML engineers, Data engineers to build and deliver multiple machine learning applications and data products.
  • Live performance of model, data deep mining, total feature engineering, Active model run time scoring, Threshold improvement to enchance model efficacy,label generation, adhoc analysis, Presenting top rated models to teams and leadership

Assistant professor & Data scientist

Confidential, Dallas, Texas

Responsibilities:

  • Searched, organized, and analyzed clinical data regarding Liver cancer. Used statistical analysis to determine data behavior. Apache spark is used to develop multiple ML model and API keys during project.
  • Used machine learning on the data to create predictive models regarding Liver cancer outcomes such as live Blood based markers, weight and many other measurable features.
  • Discussed data and models with research and development team to make the best possible products for consumers. Designed and revised forms, manuals, and other company documents. This includes product input forms, test results forms, and other files that are for internal and external use.
  • Managed coworkers on certain projects regarding QA testing and product development. Attended scientific conferences and teleconferences to help market, promote, and demonstrate company products.

Post-doctoral Research & Data scientist

Confidential, New York

Responsibilities:

  • Data Mining, Modeling, and Algorithm Development Implemented novel algorithms for modelling chemical systems and mining chemical databases using extensive application of graph theory.
  • Algorithms include conformation sampling, drug design, and reaction based design through graph manipulations. Conceptualization and design of a hierarchical database enabling fast graph isomorphism searches for storing and retrieving chemical information.
  • Software Development Experience developing object-oriented code in a team environment using software design patterns, revision control and unit-testing. Cross-platform software development on Linux, Windows, Mac Extensive experience in graph theory applications, relational databases and multithreaded code.
  • Predictive Modeling, and Statistical Analysis Using artificial neural networks (ANN) for correlating chemical structure to drug features. Identified novel molecules with activity against a cancer-causing protein target.
  • Apache spark is used to develop ML models, API keys and various applications during projects.

Notable Projects developed as Data Scientist

Confidential

Responsibilities:

  • This is a worldwide competition organized by Microsoft Professional Program (MPP) Capstone challenge 2019. Nearly 711 data scientists competed in this challenging project. My mortgage prediction model accuracy is 77 percentile and ranked 6th in the competition. Apache spark is used to develop ML models, API keys and various applications during projects.
  • Executed and Performed feature engineering, data visualization, missing data and variables.
  • Built multiple ML models performing Linear regression, Decision Forest Regression, Boosted forest regression, Neuronal Network Regression algorithms, Baines regression.
  • Evaluated model on multiple evaluation metrics like Top-K Mean Average Precision, KAPPA score, Precision & Recall. Models achieved 77% accuracy on test set.

We'd love your feedback!