Data Scientist Resume
Bellevue, WA
SUMMARY:
- 10+ years of experience in scientific research, in mathematical modeling, and in modeling, analysizing, and visualizing high dimensional data. Excellent experience in predictive modeling, clustering, classification, supervised/unsupervised learning, regression, Natural Language Processing, Text mining, feature engineering and in utlizing Hadoop/Pyspark ecosystem.
- Developped an algorithm to match billions of transactions records against DnB names to standardise Confidential customers‘ money transfer through ACH, Wires and commercial cards to prevent frauds
- Published 4 first - author papers in leading scientific journals which were cited more than 120 times; Contributed to three other papers, and demontrated an ability to work as a team member and independetly.
- Completed data science competitions at Confidential in the top 1%; used deep learning algorithms, scikit-learn, and Text mining tools: NLTK, word2vec, Spacy, gensim.
- Developed softwares using third-party applications and integrated them into multichannel order management systems.
OBJECTIVE:
I am interested in working on data and data-driven systems using advanced and rigorous machine-learning algorithms.
TECHNICAL SKILLS:
Advanced: Python, C++, C#, Hadoop, Pyspark, Scala, SQL, Hive, Impala, Hue, xml, AWS
Intermediate: Java Script, R, matlab, SAS, Data visualization and analyzing, text mining, supervised and unsupervised machine learning, Deep learning, Natural Language Processing, Model stacking, blending, Scikit-learn, Pandas, Numpy, Seaborn, Xgboost, lightlgb, Keras, TensorFlow, Scipy, Matplotlib, iPython notebook, Scipy
WORK EXPERIENCE:
Data Scientist
Confidential, Bellevue, WA
Responsibilities:
- Developped time series analysis to understand what attributes of campaigns work best to help build future marketing strategy
- Analyzed campaign data to understand customer behaviour and to design better campaign strategy
- Used Python, R, Oracle, seaborn, matplotlib
Confidential, San Francisco, CA
Responsibilities:
- Developped an algorithm to match billions of transactions records against DnB names to standardise Confidential customers‘ money transfer through ACH, Wires and commercial cards; led and managed the project.
- Prepared technical documents and made sure that they met the client’s expectation.
- Used Hadoop, pyspark, scikit-learn, Scipy, keras, Impala, Hive, Hue, Genism, Corpus
Software Engineer
Confidential, Hurlock, MD
Responsibilities:
- Developed and implemented custom C# software solutions into NET 4.6 Framework.
- Developed applications using Amazon, Shopify, PayPal, google map APIs and integrated them into multi-channel order management systems (MOM) and ABECAS insight.
- Researched best technologies and optimized internal order management algorithms, developed automated systems to generate reports, and designed projects.
- Developed user-friendly APIs to monitor Stocks and automated routine tasks
- Provided supports to our clients and our software users.
- Used C#, C++, Python, SQL, MongoDB, git, ITextSharp, json, xml, html, and many other packages.
Confidential
Responsibilities:
- Built a model to reduce the time a car spends in a test bench; Used pandas, numpy, Xgboost, Light Confidential, Random forest regression, Extra tree regression, Keras, Stack-net
- Stacked 12 classifiers of different signatures and selected the best model using out of fold prediction; Blended models and optimized weights.
Confidential
Responsibilities:
- Predicted financial movements using machine-learning algorithms: Used pandas, numpy, Xgboost, Light Confidential, Linear regression, Support vector machine
- Predicted what customers will re-order based on their previous order histories; Used pandas, numpy, Xgboost, Light Confidential, word2vec, NLTK
- Developed machine-learning algorithms and engineered features to identify duplicate questions; Used deep learning algorithms, Xgboost, scikit-learn, and Text mining tools: NLTK, word2vec, Spacy, gensim;
- Developed a machine-learning model to predict the number of inquiries a new listing receives based on the listing’s creation date and other features to better handle fraud control, identify potential listing quality issues, and allow owners and agents to better understand renters’ needs and preferences. I have used ensemble and tree-based models such us Xgboost, Random Forest Classifier, Extra Tree classifier, Gradient Boosting Classifier and label encodings.
- Develop algorithms which use a broad spectrum of features to predict realty prices which will allow Sberbank to provide more certainty to their customers in an uncertain economy. I have used linear regression and deep learning models after cleaning up a mess data.
Postdoctoral Research Fellow
Confidential
Responsibilities:
- Used Vector NTI and other sequence analyses software to predict optimal expression of a vector.
- Engineered and cloned DNA sequences into bacterial cells and transfused them into a mouse.
- Designed and conducted experiments, mentored undergraduate students and wrote a grant
- Analyzed large volume of ion channel data, and presented the result at scientific conferences
Research Assistant
Confidential
Responsibilities:
- Developed a generic algorithm to solve Ising model in n-dimension using a scalar field theory; Improved its efficiency by 25% using Hybrid Monte Carlo method.
- Used patch-clamp software to analyze single channel recordings of a Biological membrane.
- Extensively used C++, Python, OriginLab, and axon softwares.
- Developed an algorithm to solve Confidential equation in 2D and pattern formation of the universe.
Confidential
Responsibilities:
- Solved Fokker-Planck and Langevin equations analytically and numerically.
- Developed a model to study Brownian particles in a ratchet potential.
