We provide IT Staff Augmentation Services!

Senior Data Scientist Resume

3.00/5 (Submit Your Rating)

Atlanta, GA

SUMMARY

  • 6+ years of experience in Data Science, Data Modeling, Data Analysis, Data Warehousing, Machine Learning, Data mining with large data sets of Structured and Unstructured data, Data Acquisition, Data Validation, Predictive modeling, Statistical modeling, Data modeling, Data Visualization, Web Crawling, Web Scraping.
  • Adept in statistical programming languages like R and Python, SAS, Apache Spark including Big Data technologies like Hadoop, Hive, Pig.
  • Expertise in managing entire data science project life cycle and actively involved in all the phases of project life cycle including data acquisition, data cleaning, data engineering, features scaling, features engineering, statistical modeling (decision trees, regression models, neural networks, SVM, clustering), dimensionality reduction using Principal Component Analysis and Factor Analysis, testing and validation using ROC plot, K - fold cross-validation and data visualization.
  • Extensive Experience using machine learning models such as random forest, KNN, SVM, logistic regression and used packages such as ggplot, dplyr, lm, rpart, Random Forest, nnet, NumPy, sci-kit learn, pandas, etc., in R, SAS and python.
  • Proficient in Statistical Modeling and Machine Learning techniques (Linear, Logistics, Decision Trees, Random Forest, SVM, K-Nearest Neighbors, Bayesian, XGBoost) in Forecasting/ Predictive Analytics, Segmentation methodologies, Regression-based models, Hypothesis testing, Factor analysis/ PCA, Ensembles.
  • Deep understanding of the Data warehousing SDLC and architecture of ETL, reporting and BI tools. Excellent experience in Extract, Transfer and Load process using ETL tools likeDataStage, Informatica,DataIntegrator and SSIS forDatamigration andData Warehousing projects.
  • Expertise inDataAnalysis,DataMigration,Data Profiling, DataCleansing, Transformation, Integration, DataImport, andDataExport using multiple ETL tools such as Informatica Power Center.
  • Expertise in all aspects of Software Development Lifecycle (SDLC) from requirement analysis, Design, Development Coding, Testing, Implementation, and Maintenance.

TECHNICAL SKILLS

Languages: SQL, PL/SQL, T-SQL, Python, R

ML Algorithms: Classifications, Regression, Clustering, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm, Time Series, Support Vector Machines.

Databases: MS SQL Server, Oracle, Spark SQL, Azure SQL, AWS RDS, MS Access, HDFS, HBase, Teradata, MongoDB, Cassandra

ETL Tools: Informatica PowerCenter, SSIS, Azure Data Factory

Data Visualization: Power BI, Tableau, R shiny, Socrata

Data Modeling: Erwin, Microsoft Visio, MYSQL Workbench, ER/STUDIO

Analytics Skills: Data Cleansing, Data Visualization, Natural Language Processing, Supervised ML, Unsupervised ML, Deep Learning, Text Analytics

Cloud Technologies: AWS, Azure

Others: Toad, Eclipse, Databricks, Deep Learning, Text Mining, C, C++, Java, JavaScript, ASP, Shell Scripting, Scala NPL, SAS.

Operating Systems: Windows, Linux/Unix

PROFESSIONAL EXPERIENCE

Senior Data Scientist

Confidential - Atlanta, GA

Responsibilities:

  • Serving as lead analyst and subject matter expert of Databricks, R, Microsoft Azure, and Power BI.
  • Conduct research and make recommendations on products, services, and standards in support of procurement and development efforts of the company’s data initiatives.
  • Implementmachinelearningalgorithms for document recommendation in enterprise taking advantage of several data sources available in the enterprise.Create and validate machine learning models with Azure Machine Learning.
  • Designing a machine learning pipeline using Microsoft Azure Machine Learning to predict and prescribe and implemented a machine learning scenario for a given data problem.
  • Used R, SQL to create Statistical algorithms involving Multivariate Regression, Linear Regression, Logistic Regression, PCA, Random Forest models, Decision trees, Support Vector Machine for estimating the risks.
  • Let the implementation of new statistical algorithms and operators on Hadoop and SQL platforms and utilized optimizations techniques, linear regressions, K-means clustering, Native Bayes, and other approaches.
  • Working knowledge of Azure Fabric, Microservices, IoT & Docker containers in Azure. Azure infrastructure management & PaaS Solution Architect - (Azure AD, Licenses, Office365, DR on cloud using Azure Recovery Vault, Azure Web Roles, Worker Roles, SQL Azure, Azure Storage)
  • Used PCA and other feature engineering, feature normalization and label encoding Scikit-learn preprocessing techniques to reduce the high dimensional data(>150 features).
  • Experimented with predictive models including Logistic Regression, Support Vector Machine (SVM), Random Forest provided by Scikit-learn, XGBoost, LightGBM and Neural network by Keras to predict showing probability and visiting counts.
  • Coach, lead and mentor junior data scientists.

Environment: Python (Scikit-Learn, NumPy, Pandas, Matplotlib), R, Machine Learning algorithms, Tableau, Azure ML, PowerBI, Databricks, ADF, Data Lake, logistic regression, SAP, random forest, OLAP, MongoDB, Files, XML, Tableau.

Data Scientist

Confidential - Herndon, VA

Responsibilities:

  • Worked with business users to gather requirements and create a data flow, process flows, and functional specification documents.
  • Developed Data Mapping, Data Governance and transformation and cleansing rules for the Master Data Management Architecture involving OLTP, ODS.
  • Developed, enhanced, and maintained Snowflake Schema withindatawarehouse and datamart with conceptual datamodels.
  • Involved in extensive Data validation using SQL queries and back-end testing and used SQL for Querying the database in UNIX environment
  • Involved withDataAnalysis primarily IdentifyingDataSets, SourceData, Source MetaData, Data Definitions andDataFormats
  • Involved in data analysis and creating data mapping documents to capture source to target transformation rules.
  • Wrote, executed, performance tuned SQL Queries for Data Analysis & Profiling and wrote complex SQL queries using joins, subqueries, and correlated subqueries.
  • Wrote PL/SQL stored procedures, functions and packages and triggers to implement business rules into the application.
  • Used ER Studio and Visio to create 3NF and dimensional data models and published to the business users and ETL / teams.
  • Involved in Data mapping specifications to create and execute detailed system test plans. The Data mapping specifies what data will be extracted from an internal data warehouse, transformed, and sent to an external entity.
  • Creating or modifying the T-SQL queries as per the business requirements and worked on creating role playing dimensions, fact, snowflake, and star schemas.
  • Using ER Studio modeling tool, publishing of adatadictionary, review of the model and dictionary with subject matter experts and generation ofdatadefinition language.

Environment: Python (Scikit-Learn, NumPy, Pandas, Matplotlib), R, Machine Learning algorithms, Tableau, Azure ML, PowerBI, Databricks, ADF, Data Lake, logistic regression, SAP, random forest, OLAP, MongoDB, Files, XML, Tableau.

Data Analyst/Developer

Confidential - Herndon, VA.

Responsibilities:

  • Understood and articulated business requirements from user interviews and then convert requirements into technical specifications. Effectively communicated with the SMEs to gather the requirements.
  • Wrote, executed, performance tuned SQL Queries for Data Analysis & Profiling and wrote complex SQL queries using joins, subqueries, and correlated subqueries.
  • Worked with business users to gather requirements and create a data flow, process flows, and functional specification documents.
  • Develop and keep current, a high-level data strategy that fits with the Data Warehouse Standards and the overall strategy of the Company.
  • Worked in setting upthe SQL Server configuration settingsto resolve various resource allocation & memory issues for SQL Server databases and to setupideal memory, min/max server options
  • Developed and created the new database objects including tables,views, index, stored procedures, functions, triggers, advanced queries,and updated statistics.
  • Implemented T-SQL features like Data Partitioning, Error handling and Snapshot Isolation.
  • Designed and implemented comprehensive Backup plan and disaster recovery strategies Implemented and Scheduled Replication process for updating our parallel servers.
  • Implemented database Maintenance plans, scheduled automated SQL Server jobs, Created Alerts, Operators, Notifications, and configured database Mail.
  • Hands on experience monitoring and performance tuning using Tuning Advisor, SQL server Profiler, Activity monitor, Windows performance monitor, DBCC, DMVs, Stored procedures.

We'd love your feedback!