Data Scientist Resume
New York, NY
PROFESSIONAL SUMMARY:
- Highly efficient Data Scientist/DataEngineer with over 7+ years of experience in areas including Data Analyst, Statistical Analysis, Machine Learning, Data mining with large data sets of structured and unstructured data in banking, travel services, strong functional knowledge, business processes and latest market trends and manufactory industries.
- Actively participated in all phases of the project life cycle including data acquisition (Web Scraping), data cleaning, Data Engineering (dimensionality reduction (PCA & LDA), normalization, weight of evidence, information value), feature selection, features scaling & features engineering, Statistical modeling (decision trees, regression models, neural networks, SVM, clustering), testing and validation (ROC plot, k - fold cross validation), Association Rule Learning, Reinforcement Learning, Deep Learning and data visualization.
- Experience working full data insight cycle - from discussions with business, understanding business logic and business drivers, Exploratory Data Analysis, identifying predictors, enriching data, working with missing values, exploring data dynamics, meaning or building predictive data models.
- Proficient in Machine learning algorithm like Linear Regression, Ridge, Lasso, ElasticNet Regression, Decision Tree, Random Forests and more advanced algorithms like ANN, CNN, RNN, Ensemble methods like Bagging, Boosting, Stacking.
- Experience in integrating data, profiling, validating and data cleansing transformation and data visualization using R, SAS and Python.
- Extensive working experience with Python including Scikit-learn, Pandas and Numpy.
- Extensive experience in Hive, Sqoop, Flume, Hue and Oozie.
- Integration Architect & Data Scientist experience in Analytics, Big Data, BPM, SOA, ETL and Cloud technologies.
- Proficient in Predictive Modeling, Data Mining Methods, Factor Analysis, ANOVA, Hypothetical testing, normal distribution and other advanced statistical and econometric techniques.
- Proficient in SAS/BASE, SAS/EG, SAS/SQL, SAS MACRO, SAS/ACCESS
- Experience in end-to-end implementation of data ware house project based on the SASEG.
- Experience in extract data from database such as DB2, Oracle, and SME-IM and UNIXserver using SAS.
- Experienced the full software life cycle in SDLC, Agile and Scrum methodologies.
- Skilled in Advanced Regression Modeling, Correlation, Multivariate Analysis, Model Building, Business Intelligence tools and application of Statistical Concepts.
- Experience in DataAnalysis, DataProfiling, DataIntegration, Migration, Datagovernance and Metadata Management, Master Data Management and Configuration Management.
- Experience in designing stunning visualizations using Tableau software and publishing and presenting dashboards, Storyline on web and desktop platforms.
- Hands on experience in implementing LDA, Naive Bayes and skilled in Random Forests, Decision Trees, Linear and Logistic Regression, SVM, Clustering, neural networks, Principle Component Analysis and good knowledge on Recommender Systems.
- Proficient in Statistical Modeling and MachineLearning techniques (Linear, Logistics, Decision Trees, Random Forest, SVM, K-NearestNeighbors, Bayesian, XG Boost) in Forecasting/ Predictive Analytics, Segmentation methodologies, Regression based models, Hypothesis testing, ANOVA, time series analysis, Factor analysis/ PCA, Ensembles.
- Experienced in Big Data with Hadoop 2, HDFS, MapReduce, and Spark.
- Experienced in Spark 2.1, Spark SQL and PySpark.
- Skilled in using dplyr and pandas in R and python for performing exploratory data analysis.
- Experience working with data modeling tools like Erwin, PowerDesigner and ERStudio.
- Experience in designing star schema, Snow flake schema for DataWarehouse, ODS architecture.
- Experience in designing stunning visualizations using Tableau software and publishing and presenting dashboards, Storyline on web and desktop platforms.
- Good understanding of TeradataSQLAssistant, TeradataAdministrator and data loadExperience with Data Analytics, Data Reporting, Ad-hoc Reporting, Graphs, Scales, Pivot Tables and OLAP reporting.
- Having good experience in NLP with Apache, Hadoop and Python.
- Extensive experience in Text Analytics, developing different Statistical MachineLearning, Data Mining solutions to various business problems and generating data visualizations using R, Python and Tableau.
TECHNICAL SKILLS:
Languages: HTML5, css3, XML, C, C++, DHTML, WSDL, R/R Studio, SAS Enterprise Guide, SAS R, R (Caret, Weka, ggplot), Perl, MATLAB, Mathematica, FORTRAN, DTD, Schemas, Json, Ajax, Java, Scala, Python (NumPy, SciPy, Pandas, Genism, Keras), SQL, PL/SQL, Pig Latin, HiveQL, Java Script, Shell Scripting
Databases: Microsoft SQL Server 2008,10,12,14,16, MySQL 4.x/5.x, Oracle 10g, 11g, 12c, DB2, Teradata, Netezza
NO SQL Databases: HBase, Cassandra, MongoDB, MariaDB
Business Intelligence Tools: Tableau, Tableau server, Tableau Reader, Splunk, SAP Business Objects, OBIEE, QlikView, SAP Business Intelligence, Amazon Redshift, or Azure Data Warehouse
Development and Cloud Computing Tools: Microsoft SQL Studio, Eclipse, NetBeans, IntelliJ, Amazon AWS, Azure
Development Methodologies: Agile/Scrum, Waterfall, UML, Design Patterns
Version Control Tools and Testing/ ETL Tools: API Git, SVM, GitHub, SVN and JUNIT, Informatica Power Centre, SSIS.
Reporting Tools: MS Office (Word/Excel/Power Point/ Visio/Outlook), Crystal reports XI, SSRS, Cognos7.0/6.0.
Data Modelling Tools: Erwin r 9.6, 9.5, 9.1, 8.x, Rational Rose, ER/Studio, MS Visio, Oracle Designer, SAP Power designer, Enterprise Architect.
Operating Systems: All versions of UNIX, Windows, LINUX, Macintosh HD
PROFESSIONAL EXPERIENCE:
Confidential, New York, NY
Data Scientist
Responsibilities:
- Worked as a Data Modeler/Analyst to generate Data Models using Erwin and developed relational database system.
- Used R, Python, MATLAB and Spark to develop a variety of models and algorithms for analytic purposes.
- Worked with Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames, and Pair RDD's.
- Created SQL tables with referential integrity and developed queries using SQL, SQL*PLUS and PL/SQL Designed both 3NF data models for ODS, OLTP systems and dimensional data models using Star and Snowflake Schemas.
- Data Warehouse, Data Migration Application under Extensive hands on experience using ETL tools like Talend, Informatica.
- Performed data integrity checks, data cleaning, exploratory analysis and feature engineer using R and Python.
- Worked as Data Architects and IT Architects to understand the movement of data and its storage and ER Studio 9.7
- A highly immersive Data Science program involving Data Manipulation and Visualization, Web Scraping, Machine Learning, GIT, SQL, Unix Commands, Python programming, No SQL, Mongo DB, Hadoop.
- Extensively worked on Data Modeling tools Erwin Data Modeler to design the data models.
- Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
- A highly immersive DataScience program involving Data Manipulation & Visualization, Web Scraping, MachineLearning, SQL, GIT, Unix Commands, No SQL, Mongo DB.
- Built analytical data pipelines to port data in and out of Hadoop/HDFS from structured and unstructured sources and designed and implemented system architecture for AmazonEC2 based cloud-hosted solution for client.
- Performed K-means clustering, Multivariate analysis and Support Vector Machines in Pythonand R.
- Professional Tableau user (Desktop, Online, and Server), Experience with Keras and Tensor Flow.
- Created map reduce running over HDFS for data mining and analysis using R and Loading & Storage data to PigScript and R for MapReduce operations and created various types of data visualizations using R, and Tableau.
- Worked on machine learning on large size data using Spark and MapReduce.
- Performed data analysis by using Hive to retrieve the data from Hadoopcluster, SQL to retrieve data from Oracle database.
- Developed Spark/Scala, Python for regular expression (regex) project in the Hadoop/Hive environment with Linux/Windows for big data resources.
- Designed tables and implemented the naming conventions for Logical and Physical DataModels in Erwin 7.0.
- Designed logical and physical data models for multiple OLTP and Analytic applications.
- Performance tuning of the database, which includes indexes, and optimizing SQL statements, monitoring the server.
- Wrote simple and advanced SQL queries and scripts to create standard and Adhoc reports for senior managers.
- Created S3 buckets and managed roles and policies for S3 buckets. Utilized S3 buckets and Glacier for file storage and backup on AWS cloud. Used DynamoDB to store the data for metrics and backend reports.
- Worked with ElasticBeanstalk for quick deployment of services such as EC2 instances, Load balancer, and databases on the RDS on the AWS cloud environment.
- Used Java code to connect AWSS3 buckets by using AWS SDK, to access media files related to the application.
- Created Data Quality Scripts using SQL and Hive to validate successful data load and quality of the data. Created various types of data visualizations using Python and Tableau.
- Used Amazon Simple Workflow service (SWF) for data migration in data centers which automates the process and tracks every step and logs are maintained in S3 bucket.
- Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python.
- Programmed a utility in Python that used multiple packages (numpy, scipy, pandas)
- Implemented Classification using supervised algorithms like Logistic Regression, Decision trees, Naive Bayes, KNN.
- As Architect delivered various complex OLAP databases/cubes, scorecards, dashboards and reports.
- Updated Pythonscripts to match data with our database stored in AWS Cloud Search, so that we would be able to assign each document a response label for further classification.
- Used Teradata utilities such as Fast Export, MLOAD for handling various tasks data migration/ETL from OLTP Source Systems to OLAP Target Systems
- Created SSIS Packages using Pivot Transformation, Execute SQL Task, Data Flow Task, etc. to import data into the data warehouse.
- Developed and implemented SSIS, SSRS and SSAS application solutions for various business units across the organization.
Environment: Python, SQL, GIT, HDFS, Pig, Hive, Oracle, DB2, Tableau Unix Commands, NoSQL, MongoDB, SSIS, SSRS, SSAS, AWS,S3,EC2,RDS,SWF,Dynamo DB, Glacier, Erwin, Tableau, OBIEE.
Confidential, Orangeburg, NY
Data Scientist
Responsibilities:
- Performed Data Profiling to learn about behavior with various features such as traffic pattern, location, time, Date and Time etc.
- Application of various machine learning algorithms and statistical modeling like decision trees , regression models and K-Means using Python and R .
- Developed clinical NLP methods that ingest large unstructured clinical data sets, separate signal from noise, and provide personalized insights at the patient level that directly improve our analytics platform.
- Used NLP methods for information extraction, topic modeling, parsing, and relationship extraction.
- Worked with NLTK library for NLP data processing and finding the patterns.
- Used clustering technique K-Means to identify outliers and to classify unlabeled data.
- Ensured that the model has low False Positive Rate.
- Created and designed reports that will use gathered metrics to infer and draw logical conclusions of past and future behavior.
- Worked on Natural Language Processing with NLTK module of python for application development and automated customer response.
- Utilized statistical Natural Language Processing for sentiment analysis, mine unstructured data, and create insights.
- Worked on feature engineering such as feature creating, feature scaling and One-Hot encoding with Scikit-learn .
- Performed Logistic Regression, Random forest, Decision Tree, SVM to classify package is going to deliver on time for the new route.
- Implemented public segmentation by implementing k-means algorithm.
- Implemented rule-based expertise system from the results of exploratory analysis and information gathered from the people from different departments.
- Generated detailed report after validating the graphs using R, and adjusting the variables to fit the model.
- Performed Data Cleaning, features scaling, features engineering using pandas and NumPy packages in python.
- Parsed the unstructured data into semi-structured format by writing complex algorithms in pyspark
- Created Data Quality Scripts using SQL and Hive to validate successful data load and quality of the data.
- Worked on Spark SQl, created data frames by loading data from Hive tables and created prep data and stored in AWS S3
- Written MapReduce code to process and parsing the data from various sources and storing parsed data into HBase and Hive using HBase - Hive Integration.
- Created SQL tables with referential integrity and developed advanced queries using stored procedures and functions using SQL server management studio.
- Experienced in building and migrating the complex ETL pipelines from on premise system to cloud database such as Azure SQL Data warehouse , Spark to make the system grow elastically
- Expertise in data modeling on Azure SQL database and SQL Datawarehouse and building Azure Data Catalogue
- Worked on Cassandra Data modelling, NoSQL Architecture , DSE Cassandra Database administration. Key space creation, Table creation, Secondary and Solr index creation, User creation & access administration
- Extensively used the Teradata Stand-alone utilities Fast load/Multiload/Fast Export/TPT to Load/Export data into/from database objects
- Wrote transformations for data conversions into required form based on the client requirement using Teradata ETL processes
- Used packages like Dplyr, tidyr& ggplot2 in R Studio for Data visualization and generated scatter plot and high low graph to identify relation between different variables.
- Created various types of data visualizations using Python and Tableau.
- Communicated the results with operations team for taking best decisions.
- Collected data needs and requirements by Interacting with the other departments.
- Assisted with compilation of business glossary metadata, data quality rules, data lineage, data mappings/work flows, metrics, thresholds, and verification in DOORS and Informatica Metadata Manager.
Environment: R, Python 2.x,Linux, Spark, Tableau Desktop, SQL Server 2012, Microsoft Excel, Spark SQL, PySpark, Teradata.
Confidential, New Haven, CT
Data Scientist
Responsibilities:
- Developed applications of MachineLearning, Statistical Analysis, and Data Visualizations with challenging data Processing problems in sustainability and biomedical domain.
- Goal is to identify the subtypes in autism for the development of targeted and more effective therapies.
- We used hierarchical clustering methods to identify the clusters in the data based on some important features, further analysis to identify the most significant brain volumes is under way.
- Designed and developed Natural Language Processing models for sentiment analysis.
- Worked on Natural Language Processing with NLTK module of python for application development for automated customer response.
- Applied concepts of probability, distribution and statistical inference on given dataset to unearth interesting findings through the use of comparison, T-test, F-test, R-squared, P-value etc.
- Applied linear regression, multiple regression, ordinary least square method, mean-variance, the theory of large numbers, logistic regression, dummy variable, residuals, Poisson distribution, Bayes, Naive Bayes, fitting function etc to data with help of Scikit, SciPy, NumPy and Pandas module of Python.
- Applied clustering algorithms i.e. Hierarchical, K-means with help of Scikit and SciPy.
- Developed visualizations and dashboards using ggplot, Tableau
- Worked on development of data warehouse, DataLake and ETL systems using relational and non-relational tools like SQL, No SQL.
- Built and analyzed datasets using R, SAS, MATLAB, and Python (in decreasing order of usage)
- Applied linear regression in Python and SAS to understand the relationship between different attributes of the dataset and causal relationship between them
- Performs complex pattern recognition of financial time series data and forecast of returns through the ARMA and ARIMA models and exponential smoothening for multivariate time series data
- Used ClouderaHadoopYARN to perform analytics on data in Hive.
- Wrote Hive queries for data analysis to meet the business requirements.
- Expertise in BusinessIntelligence and data visualization using R and Tableau.
- Expert in Agile and ScrumProcess.
- Validated the Macro-Economic data (e.g. BlackRock, Moody's etc.) and predictive analysis of world markets using key indicators in Python and machine learning concepts like regression, Bootstrap Aggregation and RandomForest.
- Worked in large-scale database environments like Hadoop and MapReduce, with working mechanism of Hadoop clusters, nodes and Hadoop Distributed File System (HDFS)
Environment: AWS, MS Azure, Cassandra, Spark, HDFS, Hive, Pig, Linux, Python (Scikit-Learn/SciPy/NumPy/Pandas), R, SAS, SPSS, MySQL, Eclipse, PL/SQL, SQL connector.
Confidential, Charlotte, NC
Data Analyst
Responsibilities:
- Developed Internet traffic scoring platform for ad networks, advertisers and publishers (rule engine, site scoring, keyword scoring, lift measurement, linkage analysis).
- Developed ETL processes for data conversions and construction of data warehouse using IBM InfoSphere DataStage.
- Collaborated with database engineers to implement ETL process, wrote and optimized SQL queries to perform data extraction and merging from SQL server database.
- Responsible for Data Cleaning, features scaling, features engineering by using NumPy and Pandas in Python.
- Performed Exploratory Data analysis (EDA) to maximize insight in to the dataset, detect the outliners and extract important variables by graphically and Numerically.
- Applied resampling methods like Synthetic Minority Over Sampling Technique (SMOTE) to balance the classes in large data sets.
- Implemented algorithms such as Principal Component Analysis (PCA) and t-Stochastics Neighborhood Embedding (t-SNE) for dimensionality reduction and normalize the large datasets.
- Implemented classification algorithms such as Logistic Regression, K-NN neighbors and Random Forests to predict the Customer credit history and Payments activity on historical data to get predicted label whether the customers are eligible for Credit card offers and report the marketing team.
- Built machine learning models and constructed multilayer perceptions for Deep Neural Networks (DNN) to identify fraudulent applications for loan pre-approvals and to identify fraudulent credit card transactions using the history of customer transactions and compared the results.
- Perform model tuning to find the best hyperparameter fit for the algorithm to achieve better results
- Ensemble methods were used to increase the accuracy of the model with different Bagging and Boosting methods.
- Performed data visualization and Designed dashboards with Tableau, and generated complex reports, including charts, summaries, and graphs to interpret the findings to the team and stakeholders. measured the performance using Confusion matrix and Classification report.
- Evaluated models using Cross Validation, Log loss function, ROC Curves and AUC for feature selection.
Environment: NumPy, Pandas, Matplotlib, Seaborn, Scikit-Learn, Tableau, SQL, Linux, Git, Microsoft Excel, Random Forests, SVM, t-SNE, PCA, Tensor Flow, Python, DNN, K-NN
Confidential
Data Analyst
Responsibilities:
- Collaborated with database engineers to implement ETL process, wrote and optimized SQL queries to perform data extraction and merging from SQL server database.
- Gathered, analyzed, and translated business requirements, communicated with other departments to collected client business requirements and access available data.
- Responsible for Data Cleaning, features scaling, features engineering by using NumPy and Pandas in Python .
- Conducted Exploratory Data Analysis using Python Matplotlib and Seaborn to identify underlying patterns and correlation between features.
- Used information value, principal components analysis, and Chi square feature selection techniques to identify.
- Applied resampling methods like Synthetic Minority Over Sampling Technique ( SMOTE) to balance the classes in large data sets.
- Designed and implemented customized Linear regression model to predict the sales utilizing diverse sources of data to predict demand, risk and price elasticity.
- Experimented with multiple classification algorithms, such as Logistic Regression , Support Vector Machine (SVM) , Random Forest , AdA boost and Gradient boosting using Python Scikit-Learn and evaluated the performance on customer discount optimization on millions of customers.
- Used F-Score, AUC/ROC , Confusion Matrix and RMSE to evaluate different model performance
- Performed data visualization and Designed dashboards with Tableau , and generated complex reports, including charts, summaries, and graphs to interpret the findings to the team and stakeholders.
- Used Keras for implementation and trained using cyclic learning rate schedule.
- Overfitting issues was resolved by batch norm, dropout helped to overcome the issue.
- Conducted in-depth analysis and predictive modelling to uncover hidden opportunities; communicate insights to the product, sales and marketing teams.
- Built models using Python and Pyspark to predict the probability of attendance for various campaigns and events.
Environment: : NumPy, Pandas, Matplotlib, Seaborn, Scikit-Learn, Tableau, SQL, Linux, Git, Microsoft Excel, PySpark-ML, Random Forests, SVM, Tensor Flow, Keras
Confidential
Jr. Data Analyst
Responsibilities:
- Built time series models with ARIMA in R to make budget forecasting
- Developed risk assessment models by using Decision Trees and Analytic Hierarchy Process
- Designed and maintained comprehensive dashboards and metrics to enable real-time business decisions
- Coded SQL queries to extract data and identify granularity issues and relationships between datasets and recommended solutions
- Involved in manipulating, cleansing & processing of data using Excel, Access and SQL
- Compared the source data with historical data to perform statistical analysis
- Performed data preprocessing and data cleaning, collected and organized data
Environment: : MS Access, Excel, R, SQl, ETL
