We provide IT Staff Augmentation Services!

Data Scientist Resume

3.00/5 (Submit Your Rating)

DallaS

SUMMARY:

  • Professional qualified Data Scientist withover 05 yearsof experience onData Science and Analyticsin Banking, Insurance and Telecom Domain.
  • Designing and developing various machine learning frameworks using python, R, and Matlab.
  • Collaborated with data engineers to implementETLprocess, wrote and optimizedSQLqueries to perform data extraction and merging fromOracle 12c.
  • Rich Experience in managing entiredata science project life cycleand involved in all phases, includingdata extraction, data cleaning, statistical modelinganddata visualization,with large datasets ofstructuredandunstructured data.
  • Hands - on experience inMachine Learningalgorithms such asLinear Regression, GLM, CART, SVM, KNN, LDA/QDA, Naive Bayes, Random Forest, SVM, Boosting, K-means ClusteringHierarchical clustering, PCA, Feature Selection, Collaborative Filtering, Neural Networks andNLP
  • Professional working experience withPython 2.X / 3.Xlibraries includingMatplotLib, Numpy, Scipy, Pandas, Beautiful Soup, Seaborn, Scikit-learnandNLTKfor analysis purpose.
  • Experience in implementing data analysis with various analytic tools, such asAnaconda 4.0 / 2.X (Jupyter Notebook, Spyder), R 2.15 / 3.0 (Reshape, ggplot2, Dlpr, Car, Mass and Lme4)SAS 9.3, Matlab 8.0andExcel 2010/2013.
  • Experience withdata visualizationsusingPython 2.X / 3.XandR 2.15 / 3.0and generating dashboard withTableau 8.0 / 9.2 / 10.0.
  • Working experience inStatistical Analysis and TestingincludingHypothesis test, Anova, Survival Analysis, Longitudinal Analysis, Experiment Design and Sample Determination andA/B test.
  • Hands-on experience in importing and exporting data usingRelational DatabaseincludingOracle 11g / 12c, MySQL 5.0andMS SQL Server, andNoSQL databaselikeMongoDB 3.3 / 3.4.
  • Working experience in big data environmentlikeHadoop Ecosystem 1.X / 2.XincludingHDFS, MapReduce, Hive 0.11, HBase 0.9,Spark Framework 1.4 / 1.6 / 2.0includingPysparkMLlibandSparkSQL
  • Working experience inversion controltools such asGit 2.Xto coordinate work on file with multiple team members.
  • Employing variousSDLCmethodologies such asAgileandSCRUMmethodologies.
  • Good team player and quick-learner; highly self-motivated person with good communication and interpersonal skills.

TECHNICAL PROFICIENCY:

Machine Learning Algorithms: Analytic Tools Linear regression, SVM, KNN, Naive Bayes,Anaconda 4.0 / 2.X (Jupyter NotebookLogistic regression, LDA/QDA, SVM, CART,Spyder), R 2.15 / 3.0 (Reshape, ggplot2Random Forest, Boosting, K-means clustering, Dlpr, Car, Mass and Lme4), SAS 9.3, Matlab Hierarchical clustering, Collaborative filtering,8.0, Mathematica 9.0, Excel Neural Network, NLP.

Statistical Analysis Programming Language: Hypothesis Test, ANOVA, Survival Analysis,Python 2.X & 3.X (numpy, scipy, pandasLongitudinal Analysis, Experiment Design andseaborn, beautiful soup, scikit-learn, NLTK)Sample Determination, A/B TestSQL, C

Hadoop Ecosystem (1.X & 2.X): Spark Framework (1.4 & 1.6& 2.0) HDFS, MapReduce, Hive 0.11, Hbase 0.9SparkSQL, Pyspark, Mllib

Relational Database: Data Visualization MySQL 5.0, Oracle 11g / 12c, MS SQLTableau 8.0 /9.2 / 10.0 , D3.js 3.X / 4.XServer R-ggplot2, Python-Matplotlib

NoSQL: Version Control MongoDB 3.3 / 3.4Git 2.X

Operation System: Windows 7 / 10, Mac OS

Programming Languages: Python, R, Matlab, SQL, UNIX, MongoDB, Spark, Hadoop, Lua, Torch.

Machine Learning and Deep learning Techniques: Trees, Bayes Model, SVM, Ensemble Methods, Neural Networks, RNN, KNN, CNN, MLP, Ensemble SVM, Majority voting, Linear models, Classification, Regression, Clustering, Kernel methods, Memory Networks, LSTMs, Dimension reduction, Deep belief networks, Statistical tests.

Python Libraries: Scikit, pandas, Numpy, Scipy, Theano, Keras, Matplotlib, pymongo.

R Libraries: dplyr, ggplot2, jsonlite, plyr, rvest, rjson, httr, xml2, curl

PROFESSIONAL EXPERIENCE:

Confidential, Dallas

Data Scientist

Responsibilities:

  • Designing and developing various machine learning frameworks using python, R, and Matlab.
  • Collaborated with data engineers to implementETLprocess, wrote and optimizedSQLqueries to perform data extraction and merging fromOracle 12c.
  • Used MovieLense Data sets to recommend New Movies to the customer, based upon the reviews.
  • Collected unstructured data fromMongoDB 3.3and completed data aggregation.
  • Performed data integrity checks, data cleaning, exploratory analysis and feature engineer using Python 3.5.
  • Conducted analysis onassessingcustomerconsumingbehaviors and discover value of customers with RMFanalysis;applied customer segmentation with clusteringalgorithms such asK-Means ClusteringandHierarchical Clustering.
  • Developed personalized products recommendationwithMachine Learningalgorithms,including Collaborative filteringandGradient Boosting Tree,to better meet the needs of existing customersand acquire new customers.
  • UsedPython 3.X (numpy, scipy, pandas, scikit-learn, seaborn, NLTK)andSpark 1.6 / 2.0 (PySpark, MLlib)todevelop variety of models andalgorithmsfor analyticpurposes.
  • Coordinatedthe execution ofA/Bteststo measuretheeffectiveness of personalized recommendation system.
  • Performed data visualization withTableau10.0, MeteorJSand generated dashboards to present the findings.
  • Recommendedand evaluatedmarketing approachesbased on quality analyticsofcustomerconsuming behavior.
  • Determinedcustomersatisfactionand helped enhance customer experience usingNLP.
  • UsedGit 2.Xto apply version control. Tracked changes infilesand coordinatedwork on the files among multiple team members.

Environment: Python 3.X, Oracle 12c,MongoDB 3.3,Spark 1.6/ 2.0(Pyspark, MLlib, Spark SQL), Tableau 10.0, D3.js 3.X / 4.X, Git 2.X

Confidential, Wichita, Kansas

Data Scientist

Responsibilities:

  • My role was to perform statistical analysis for worldwide insurance market research data collected from surveys, give insights into specific business problems, preparing technical documentation and present results in a simple and clear fashion to a non-statistical audience.
  • Identifiedrisk leveland eligibilityof newinsurance applicants withMachine Learning algorithms.
  • Predicted the claim severityto understand future loss and ranked importance offeatures.
  • UsedPython 3.XandSpark 1.4(PySpark, MLlib)to implementdifferentmachine learning algorithmsincludingGeneralized Linear Model, SVM, Random Forest, BoostingandNeural Network.
  • Evaluated and optimized performance of models,tuned parameters withK-Fold Cross Validation.
  • Provided analytical support tounderwritingand pricingby preparing and analyzing data to be used in auctorial calculations
  • Designed dashboards withTableau9.2and MeteorJS provided complex reports, including summaries, charts, andgraphsto interpret findings to team and stakeholders.
  • Identifiedprocess improvements that significantly reduce workloads or improve quality.
  • UtilizedSQLandHiveQLto query, manipulate data from variety data sources including Oracle 10g andHDFS, while maintaining data integrity.
  • Worked on data cleaning, data preparation and feature engineering withPython3.Xincluding Numpy, Scipy, Pandas, Matplotlib, SeabornandScikit-learn.
  • Worked with the version control tools, such asGit 2.X, to keep versions attributed from different people and record project at different time points.
  • A Novel Machine Learning Framework For Phenotype Prediction Based On Genome-Wide DNA Methylation Data. Published IJCNN 2017
  • Recognizing handwritten digits using artificial neural network
  • Implemented a project using python, skit-learn, Numpy, Scipy and UNIX for detecting handwritten digits using artificial neural networks (Natural Language Processing) and convolutional neural networks.
  • Developed a database using JAVA, SQL, and UNIX for managing a library, with one user as a librarian with limited access and the other as a database administrator.

Environment: Python 3.X, Oracle 10g, Hive 0.11,HDFS,Spark 1.4, Tableau 9.2, Git 2.X

Confidential, Wichita, Kansas

Data Scientist

Responsibilities:

  • Developed WebSets for variation and Time Series analysis, of Stock Market, and for Rent updates on Sulekha.Com.
  • Prediction of Multiple Regression analysis.
  • Designed and developed an advanced recommendation system for real estate and finance customers utilizing R and Python.
  • UtilizedSQLto extract data fromSQL Server 11.0, and MongoDB, to prepare data for analysis. classified customers byRFManalysis,clusteringand regression model withR 3.0, selected customers with high value and improved their retention rate by sending ads and coupons

Environment: Excel 2013, R 3.0, Hadoop 1.X, MS SQL Server 2012, Tableau 8.1

Confidential

Data Scientist

Responsibilities:

  • UsedPython 2.7to applytime seriesmodels,clustering algorithmand other data mining methods to explore the fast growth opportunities of our clients
  • Analyzed the traffic queries of Baidu search engine usingclassificationalgorithm.
  • Assisted to improve the liquidity of our ads model.
  • Based on the data of clients and traffic, designed comprehensive analysis to optimize products and explored the strengths and weaknesses of products.
  • Boosted marketing for client by designing a cell phone data rating engine in Python, PL/SQL, and UNIX/Linux to suggest suitable cell phone plans based on data usage.
  • Increased efficiency of client customer support by implementing a call center application using R, Python and SQL that tracked customer complaint history and assigned defects accordingly.
  • Enhanced customer buying experience using Machine learning, R, python and SQL
  • Team Member/SPOC R50 Release, One Siebel
  • Ensured smooth access for customers by developing scripts in Python and Unix for managing the crash reports and log levels of Unix and Windows servers, maintaining server network configuration components using Siebel 8.1 and Oracle, analyzing and providing RCA for server defects, and developing and publishing documents for avoiding defects in future releases; set up Scrum daily between Delivery Managers, Dev Team, Test Team, and Admin Team.

Environment: Excel 2010, Python 2.7, Hadoop 1.X, Mapreduce, MS SQL Server 2008, Tableau 8.0

We'd love your feedback!