Data Scientist Resume
New York, NY
SUMMARY
- Data Scientist with 6+ years of constructive experience in analyzing and solving real world problems in the Banking/Finance, E - commerce and Healthcare domains.
- Expertise in designing Statistical and Predictive models on large-scale Structured and Unstructured Data.
- Proficiency in implementing various Supervised and Unsupervised Machine Learning algorithms.
- Highly efficient in performing Data Scraping, Data Cleansing, Feature Scaling and Feature Engineering.
- Skilled in using statistical methods including Experiment Design and Sample Determination, Hypothesis Test, ANOVA, Chi-square, A/B test, Z- test, T- test, Cross-Validation.
- Proficient in Machine Learning algorithms such as Linear and Logistic Regression, SVM, KNN, Naïve Bayes, Decision Trees, Random Forests, Gradient Boosting, K-means Clustering, Hierarchical clustering, PCA, Collaborative Filtering, Neural Networks and NLP.
- Extensive experience in statistical programming languages and data analytical tools such as Anaconda 4.x/3.x/2.x (Jupyter Notebook), R 3.x (Reshape, ggplot2, Dplr, Car, Mass, Lme4), MATLAB and SQL.
- Adept in Python 2.x/3.x libraries including Matplotlib, Numpy, Pandas, Seaborn, Scikit-learn for analysis purpose.
- Expert in designing visualizations using Tableau, Power BI, Microsoft Excel, D3.js and publishing and presenting Dashboards, Storylines on web and desktop platforms.
- Hands on experience with RDBMS including MySQL, SQL Server, Oracle.
- Experience working with key-value store, and documented-oriented NoSQL database including HBase, MongoDB and Cassandra.
- Working knowledge of Big Data Topics Spark, Hive, Hadoop and MapReduce.
- Strong experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification in both Waterfall and Agile methodologies.
- Skilled at delivery and presentation of analytical business insights, as well as planning analytic roadmaps to facilitate data-driven decision making.
- Ability to independently and effectively organize and manage multiple assignments with excellent analytical and problem-solving skills.
- A strong team player as well as capable of working independently.
TECHNICAL SKILLS
Machine Learning Algorithms Statistical Analysis:: Linear and Logistic Regression, SVM, KNN,Experiment Design and Sample DeterminationNaïve Bayes, Decision Trees, Random Forests,Hypothesis Test, ANOVA, Chi-square, A/B testGradient Boosting, K-means Clustering, Z- test, T- test, Cross-Validation Hierarchical clustering, Collaborative FilteringPCA, Neural Networks and NLP
Database Big Data:: Relational - My SQL, SQL Server, Oracle 11G,Apache Hadoop, Map Reduce, HDFS, HiveNoSQL - MongoDB, HBase, CassandraSpark, Sqoop, Yarn, Flume, Kafka
Programming Languages Data Science Libraries:: Python 2.x/3.x, R 3.x, PL/SQL, JavaPython - Numpy, Pandas, Scipy, Scikit-learn, NLTK, Seaborn, Matplotlib, Plotly, R - Reshape, ggplot2, Dplr, Car, Mass, Lme4, Shiny, Caret
BI/ Visualization tools Operating Systems:: Tableau, Power BI, D3.js, QlikViewWindows XP/7/8.1/10, Linux - Ubuntu.
Version Control Methodologies:: Git, Subversion, SCMAgile (Scrum), Waterfall
PROFESSIONAL EXPERIENCE
Confidential - NEW YORK, NY
DATA SCIENTIST
RESPONSIBILITIES:
- Extracted, transformed and merged data from multiple data stores.
- Performed Exploratory Data Analysis on transactions history using Hive on HDFS.
- Worked with NoSQL Database MongoDB to extract some of the required user data.
- Performed Data cleaning, feature scaling, feature engineering. Created a utility matrix using these new features.
- Initial models were built using supervised classification techniques like K - Nearest Neighbor (KNN), Logistic Regression and Random Forests with Principal component analysis to identify important features.
- Built models using K - means clustering to create user groups.
- Experimented with models built using content based recommendation and user based recommendation.
- Achieved scalability, accuracy and speed with the help of Collaborative Filtering.
- Created a hybrid model to support new user recommendation and existing users with changing trends.
- Fine-tuned the algorithm using regularization term to overcome the problem of over fitting.
- Used RMSE and Mean Average Precision to evaluate recommender’s performance in both stimulated environment and real world
- Deployed the model in production and monitored user activity and add-on sales from items that were recommended and not searched.
- Used the results to tune the parameters and rebuild the model.
- Created visualizations to convey results and analyze data using Tableau 9.3.
ENVIRONMENT: Python 3.6, Hive, Spark, HDFS, Mongo DB, Tableau 9.3
Confidential - NEW YORK, NY
DATA SCIENTIST
RESPONSIBILITIES:
- Used Predictive Modeling, Statistics, Machine Learning, Data Mining, and other aspects of data analytics techniques to collect, explore, and extract insights.
- Handled large amounts of structured and unstructured data.
- Actively involved in designing and developing data ingestion, aggregation, and integration in Hadoop environment.
- Developed Sqoop scripts to import export data from relational sources and handled incremental loading on the date.
- Experience in creating Hive Tables, Partitioning and Bucketing.
- Performed Data Cleaning, Feature Scaling, and Feature engineering using Python 3.6 and Python libraries like Pandas and Numpy.
- Used Jupyternotebook for writing Python scripts for / testingdatasets & making prediction.
- Implemented Naïve Bayes, Decision Trees, Random Forest and Gradient Boosting for predictive analysis using python Scikit-Learn.
- Deployed data results with data visualization tools like Python Seaborn and Matplotlib.
- Used Cross Validation to evaluate models for detecting Overfitting.
- Used A/B test and Hypothesis test to check the accuracy of the model.
- Visualized, interpreted, report findings and developed strategic uses of data using Tableau 9.3 and D3.js, and created interactive Dashboards.
- Documented all programs and procedures to ensure an accurate historical record of work completed on assigned project as well as to improve quality and efficiency.
- Sent the reports to customer managers and management to reduce churn rate and improve customer satisfaction.
- Used feedback from customer managers to tune the models and add/remove features.
- Followed Agile methodology in entire project lifecycle.
- Used Git for version control of the code.
ENVIRONMENT: Python 3.6, Anaconda 3.0, Hadoop, Hive, Sqoop, D3.js, Tableau 9.3, Agile, Git
Confidential
DATA ANALYST
RESPONSIBILITIES:
- Involved in Analysis, Design and Implementation/translation of Business User requirements.
- Worked on collection of large sets using Python scripting.
- Worked with both Structured as well as Unstructured data.
- Worked on loading the data from MySQL to NoSQL database-Cassandra where necessary.
- Performed data analysis and data profiling using complex SQL queries on SQL Server.
- Identified inconsistencies in data collected from different source.
- Worked with business owners/stakeholders to assess Risk impact, provided solution to business owners.
- Experienced in determining trends and significant data relationships Analyzing using Advanced Statistical Methods.
- Carrying out specified data processing and statistical techniques such as sampling techniques, estimation, hypothesis testing, time series, correlation and regression analysis Using R.
- Applied various data mining techniques: Linear Regression & Logistic Regression, classification, clustering.
- Took personal responsibility for meeting deadlines and delivering high quality work.
- Strived to continually improve existing methodologies, processes, and deliverable templates.
ENVIRONMENT: Python 2.7, R, SQL server, Cassandra, MS Excel, Tableau 8.3
Confidential
DATA ANALYST
RESPONSIBILITIES:
- Developed business process models using MS Visio to create case diagrams and flow diagrams to show flow of steps that are required.
- Designed, Implemented and automated modelling and analysis procedures on existing and experimentally created data.
- Created PL/SQL packages, Database Triggers and developed user procedures and prepared user manuals for the new programs.
- Created dynamic linear models to perform analysis using Python 2.7
- Used MS Excel, MS Access and SQL to write and run various queries.
- Used traceability matrix to trace the requirements of the organization.
- Analyze the data and create dashboards using Tableau.
- Reviewed the logical model with application developer, ETL teams, DBA’S and testing team to provide information about the data model and business requirements.
- Involved in the daily maintenance of the database that involved monitoring the daily run of the scripts as well as troubleshooting in the event of any errors in the entire process.
ENVIRONMENT: Python 2.7, PL/SQL, Oracle, MS office, MS Visio, Tableau 8.0.
ConfidentialPL/SQL DEVELOPER
RESPONSIBILITIES:
- Developed, tested, debugged and documented Oracle PL/SQL packages to implement the business requirements.
- Created database objects like Tables, Sequences, Views, Triggers, Procedures and Functions.
- Created indexes on tables for faster retrieval of the data to enhance data performance.
- Designed complex SQL queries using sub queries to retrieve data from the database.
- Tested coding modifications and assisted with application and system testing to minimize errors and downtime.
- Provided updates and status to the Management on customer issues.
- Extensively used Joins writing Views in the database.
- Maintained versioned code in Subversion in accordance with company policies and industry best practices.
ENVIRONMENT: Oracle 10g/11g, SQL* Loader, Toad, PL/SQL, Subversion
