We provide IT Staff Augmentation Services!

Data Scientist Resume

3.00/5 (Submit Your Rating)

KS

SUMMARY:

  • Designing and developing various machine learning frameworks using python, R, and Matlab.
  • Collaborated with data engineers to implementETLprocess, wrote and optimizedSQLqueries to perform data extraction from Cloud and merging fromOracle 12c.
  • Professional qualified Data Scientist withover 05 yearsof experience onData Science andAnalyticsin Banking, Insurance and Telecom Domain.
  • Rich Experience in managing entiredata science project life cycleand involved in all phases,includingdata extraction, data cleaning, statistical modelinganddata visualization,withlarge datasets ofstructuredandunstructured data.
  • Hands - on experience inMachine Learningalgorithms such asLinear Regression, GLM, CART,SVM, KNN, LDA/QDA, Naive Bayes, Random Forest, SVM, Boosting, K-means ClusteringHierarchical clustering, PCA, Feature Selection, Collaborative Filtering, Neural NetworksandNLP
  • Professional working experience withPython 2.X / 3.Xlibraries includingMatplotLib, Numpy,Scipy, Pandas, Beautiful Soup, Seaborn, Scikit-learnandNLTKfor analysis purpose.
  • Experience in implementing data analysis with various analytic tools, such asAnaconda 4.0 / 2.X(Jupyter Notebook, Spyder), R 2.15 / 3.0 (Reshape, ggplot2, Dlpr, Car, Mass and Lme4)SAS 9.3, Matlab 8.0andExcel 2010/2013.
  • Experience withdata visualizationsusingPython 2.X / 3.XandR 2.15 / 3.0and generatingdashboard withTableau 8.0 / 9.2 / 10.0.
  • Working experience inStatistical Analysis and TestingincludingHypothesis test, Anova,Survival Analysis, Longitudinal Analysis, Experiment Design and Sample DeterminationandA/B test.
  • Hands-on experience in importing and exporting data usingRelational DatabaseincludingOracle11g / 12c, MySQL 5.0andMS SQL Server 2008 / 2012, andNoSQL databaselikeMongoDB3.3 / 3.4.
  • Working experience in big data environmentlikeHadoop Ecosystem 1.X / 2.XincludingHDFS,MapReduce, Hive 0.11, HBase 0.9,Spark Framework 1.4 / 1.6 / 2.0includingPysparkMLlibandSparkSQL
  • Working experience inversion controltools such asGit 2.Xto coordinate work on file with multipleteam members.
  • Employing variousSDLCmethodologies such asAgileandSCRUMmethodologies.
  • Good team player and quick-learner; highly self-motivated person with good communication and

TECHNICAL SKILLS

Machine Learning Algorithms: Analytic Tools Linear regression, SVM, KNN, Naive Bayes,Anaconda 4.0 / 2.X (Jupyter NotebookLogistic regression, LDA/QDA, SVM, CART,Spyder), R 2.15 / 3.0 (Reshape, ggplot2Random Forest, Boosting, K-means clustering, Dlpr, Car, Mass and Lme4), SAS 9.3, MatlabHierarchical clustering, Collaborative filtering,8.0, Mathematica 9.0, Excel 2010 / 2013Neural Network, NLP.

Statistical Analysis Programming Language: Hypothesis Test, ANOVA, Survival Analysis,Python 2.X & 3.X (numpy, scipy, pandasLongitudinal Analysis, Experiment Design andseaborn, beautiful soup, scikit-learn, NLTK)Sample Determination, A/B TestSQL, C

Hadoop Ecosystem (1.X & 2.X): Spark Framework (1.4 & 1.6& 2.0) HDFS, MapReduce, Hive 0.11, Hbase 0.9SparkSQL, Pyspark, Mllib

Relational Database: Data Visualization MySQL 5.0, Oracle 11g / 12c, MS SQLTableau 8.0 /9.2 / 10.0, D3.js 3.X / 4.XServer 2008 / 2012R-ggplot2, Python-Matplotlib

NoSQL: Version Control MongoDB 3.3 / 3.4Git 2.X

Operation System: Windows 7 / 10, Mac OS

Programming Languages: Python, R, Matlab, SQL, UNIX, MongoDB, Spark, Hadoop, Lua, Torch, Tensorflow.

Machine Learning and Deep learning Techniques: Trees, Bayes Model, SVM, Ensemble Methods, Neural Networks, RNN, KNN, CNN, MLP, Ensemble SVM, Majority voting, Linear models, Classification, Regression, Clustering, Kernel methods, Memory Networks, LSTMs, Dimension reduction, Deep belief networks, Statistical tests.

Python Libraries: Scikit, pandas, Numpy, Scipy, Theano, Keras, Matplotlib, pymongo.

R Libraries: dplyr, ggplot2, jsonlite, plyr, rvest, rjson, httr, xml2, curl

PROFESSIONAL EXPERIENCE:

Confidential, KS

Data Scientist

Responsibilities:

  • Designing and developing various machine learning frameworks using python, R, and Matlab.
  • Collaborated with data engineers to implementETLprocess, wrote and optimizedSQLqueries to perform data extraction from Cloud and merging fromOracle 12c.
  • Collected unstructured data fromMongoDB 3.3and completed data aggregation.
  • Performed data integrity checks, data cleaning, exploratory analysis and feature engineer using R 3.4.0.
  • Conducted analysis onassessingcustomerconsumingbehaviors and discover value of customers with RMFanalysis;applied customer segmentation with clusteringalgorithms such asK-Means ClusteringandHierarchical Clustering.
  • Developed personalized products recommendationwithMachine Learningalgorithms,including Collaborative filteringandGradient Boosting Tree,to better meet the needs of existing customersand acquire new customers.
  • UsedPython 3.X (numpy, scipy, pandas, scikit-learn, seaborn, NLTK)andSpark 1.6 / 2.0 (PySpark, MLlib)todevelop variety of models andalgorithmsfor analyticpurposes.
  • Coordinatedthe execution ofA/Bteststo measuretheeffectiveness of personalized recommendation system.
  • Performed data visualization withTableau10.0, MeteorJSand generated dashboards to present the findings.
  • Recommendedand evaluatedmarketing approachesbased on quality analyticsofcustomerconsuming behavior.
  • Determinedcustomersatisfactionand helped enhance customer experience usingNLP.
  • UsedGit 2.Xto apply version control. Tracked changes infilesand coordinatedwork on the files among multiple team members.

Environment: R 3.X, Oracle 12c,MongoDB 3.3,Spark 1.6/ 2.0(Pyspark, MLlib, Spark SQL), Tableau 10.0, D3.js 3.X / 4.X, Git 2.X

Confidential

Data Scientist

Responsibilities:

  • Identifiedrisk leveland eligibilityof newinsurance applicants withMachine Learning algorithms.
  • Predicted the claim severityto understand future loss and ranked importance offeatures.
  • UsedR 3.X, R2.XandSpark 1.4(PySpark, MLlib)to implementdifferentmachine learningalgorithmsincludingGeneralized Linear Model, SVM, Random Forest, BoostingandNeuralNetwork.
  • Evaluated and optimized performance of models,tuned parameters withK-Fold Cross Validation.
  • Provided analytical support tounderwritingand pricingby preparing and analyzing data to be used inauctorial calculations
  • Designed dashboards withTableau9.2and MeteorJS provided complex reports, includingsummaries, charts, andgraphsto interpret findings to team and stakeholders.
  • Identifiedprocess improvements that significantly reduce workloads or improve quality.
  • UtilizedSQLandHiveQLto query, manipulate data from variety data sources includingOracle10g andHDFS, while maintaining data integrity.
  • Worked on data cleaning, data preparation and feature engineering withPython3.XincludingNumpy, Scipy, Pandas, Matplotlib, SeabornandScikit-learn.
  • Worked with the version control tools, such asGit 2.X, to keep versions attributed from differentpeople and record project at different time points.
  • A Novel Machine Learning Framework For Phenotype Prediction Based On Genome-Wide DNA Methylation Data. Published IJCNN 2017
  • Recognizing handwritten digits using artificial neural networkImplemented a project using python, skit-learn, Numpy, Scipy and UNIX for detecting handwritten digits using artificial neural networks (Natural Language Processing) and convolutional neural networks.
  • Library Management systemDeveloped a database using JAVA, SQL, and UNIX for managing a library, with one user as a librarian with limited access and the other as a database administrator.

Environment: R 3.X,R 2.X, Oracle 10g, Hive 0.11,HDFS,Spark 1.4, Tableau 9.2, Git 2.X

Confidential

Data Scientist

Responsibilities:

  • Prediction of Multiple Regression analysis.
  • Designed and developed an advanced recommendation system for real estate and finance customers utilizing R and Python.
  • UtilizedSQLto extract data fromSQL Server 11.0, and MongoDB, to prepare data for analysis.
  • classified customers byRFManalysis,clusteringand regression model withR 3.0, selected customers with high value and improved their retention rate by sending ads and coupons

Environment: Excel 2013, R 3.0, Hadoop 1.X, MS SQL Server 2012, Tableau 8.1

Confidential

Data Scientist

Responsibilities:

  • UsedPython 2.7to applytime seriesmodels,clustering algorithmand other data mining methods to explore the fast growth opportunities of our clients
  • Analyzed the traffic queries of Baidu search engine usingclassificationalgorithm.
  • Assisted to improve the liquidity of our ads model.
  • Based on the data of clients and traffic, designed comprehensive analysis to optimize products and explored the strengths and weaknesses of products.
  • Team Member, Rating Engine ApplicationBoosted marketing for client by designing a cell phone data rating engine in Python, PL/SQL, and UNIX/Linux to suggest suitable cell phone plans based on data usage.
  • Team Lead, Call Center ApplicationIncreased efficiency of client customer support by implementing a call center application using R, Python and SQL that tracked customer complaint history and assigned defects accordingly.
  • Team Member, Open ReachEnhanced customer buying experience using Machine learning, R, python and SQL
  • Team Member/SPOC R50 Release, One Siebel
  • Ensured smooth access for customers by developing scripts in Python and Unix for managing the crash reports and log levels of Unix and Windows servers, maintaining server network configuration components using Siebel 8.1 and Oracle, analyzing and providing RCA for server defects, and developing and publishing documents for avoiding defects in future releases; set up Scrum daily between Delivery Managers, Dev Team, Test Team, and Admin Team.

Environment:Excel 2010, Python 2.7, Hadoop 1.X, Mapreduce, MS SQL Server 2008, Tableau 8.0

We'd love your feedback!