We provide IT Staff Augmentation Services!

Data Scientist Resume

5.00/5 (Submit Your Rating)

NorfolK

SUMMARY

  • 9 years experience in Machine Learning, Data - mining with large datasets of Structured and Unstructured data, Data Acquisition, Data Validation, Predictive modeling, Data Visualization.
  • Experience in datastructure, design and analysis using MachineLearning Techniques and modules in PYTHON, R.
  • Experienced with SaaS based BI tools like Amazon QuickSight during my projects
  • Implemented machine learning algorithms on large datasets to understand hidden patterns and capture insights.
  • Good experience in system monitoring, development and support related activities for Hadoop and Hadoop Admin, Java/J2EE Technologies and Spark Streaming, Spark SQL.
  • Experience in working on Apache Hadoop ecosystem components like Map Reduce, HDFS, Hive, Pig, HBase, Flume, Sqoop, Oozie.
  • Working knowledge of optimization tools (Gurobi, Cplex, or Xpress)
  • Extensive Knowledge in implementation of NLP, Text Mining, Recommendation System, Deep learning, Scala.
  • Proficiency inspark using Scala for loading data from the local file systems like HDFS, Amazon S3, Relational and NoSQL databases usingSpark SQL and Import data into RDD.
  • Extensive experience in in-depth data analysis on different DB and Data Extraction. Strong knowledge in writing SQL Queries, sub-queries, and joins.
  • Skilled in providing analytic support including data importing/Extraction, data wrangling and data visualization.
  • Experience in creating Databases, Tables, Stored Procedure, DDL/DML Triggers, Views, User defined data types,effective functions,Cursors and Indexes.
  • Sound knowledge in Enterprise Data Warehousing and Business Intelligence, ETL (Extract, Transform, Load), dimensional/Hierarchical data modeling, Data mapping, Data Dictionaries.
  • Expert in Applying Advance MS excel and Adept in MS Excel with proficiency in VLOOKUP’s, Pivot Tables and understanding of VBA Macros.
  • Familiarity and good knowledge with data manipulation software Alteryx, SAS.
  • Strong understanding of HRMS/CRM databases/systems, to query and analyze customer data. Experience with Google Analytics, Tableau, MS EXCEL and SQL Server Reporting Services (SSRS).
  • Experience in designings tunning visualizations using Tableau software and publishing and presenting dashboards on web and desktop platforms.
  • Deploying, managing, and operating scalable, highly available, and fault tolerant systems on AWS
  • Well experienced in Normalization, De-Normalization and Standardization techniques for optimal performance in relational and dimensional database environments.
  • Strong experience in Software Development Life Cycle (SDLC) including Requirements, Specifications Analysis/Design and Testing as per the Software Development Life Cycle.
  • Excellent communication, teamoriented and interpersonal skills, QuickLearner, Exceptional Team Player and I have the ability to work independently as well.

TECHNICAL SKILLS

Operating Systems: Windows 10/7/XP, Linux, Unix

Databases: SQL, Hive, Impala, Pig, Spark SQL, Databases SQL-Server, MySQL, MS Access, HDFS, HBase, Teradata, Netezza, Mongo DB, Cassandra, SAP HANA.

Statistical: Microsoft Excel, R, Python

Business Intelligence Tools: IBM Cognos,Tableau

Machine learning: Classification, Regression, Clustering, Anomaly detection, Artificial Intelligence

BI Tools: Tableau, Tableau server, Tableau Reader, SAP Business Objects, OBIEE, QlikView, SAP Business Intelligence, Amazon Redshift, or AzureDataWarehouse

Languages: Spark, SQL, PL/SQL, Python, R, Scala, UML

NaturalLanguage Processing: Stanford CoreNLP, NLTK, Gensim, Apache Spark, SciKit-learn, Regular Expressions

Deep learning methods: Simple Neural nets, DNN, CNN, RNN, LSTM, GRU

PROFESSIONAL EXPERIENCE

Confidential, NORFOLK

Data Scientist

Responsibilities:

  • Extracthealth careclaims, provider, and enrollmentdata from Database and support of the triple aim for Betterhealth, better quality, and lower costs.
  • Identified data analytics opportunities like data requirements gathering, designs, data extraction, and analysis.
  • Analyze customer, behavior data, symptoms data, transaction data and campaign data to identify trends and patterns of data in different visualization techniques like Seaborn library in PYTHON.
  • Worked on NLP techniques such as PoS tagging, query pre-processing/re-structuring, deep/shallow parsing, etc
  • Worked on entire Development Architecture within the Data Lake
  • Utilized commercial optimization solvers (e.g. CPLEX, DASH, Gurobi) and software profiling & memory management tools)
  • Experience in Building cloud strategy, Enabling ‘As a Service’ (SAAS, PAAS and IAAS), Building cloud application migration business case and migration road map and Cloud cost benefit TCO & ROI analysis
  • Wrote simple and advanced SQL queries and scripts to create standard reports for senior managers.
  • Worked on data cleaning and ensured data quality, consistency, integrity using Pandas, Numpy
  • Worked on outliers identification with box-plot, K-means clustering using Pandas, Numpy
  • Participated in features engineering such as feature intersection generating, feature normalize and Label encoding with Scikit-learn preprocessing Modeled customers to discover untapped business opportunities.
  • Optimized efficiency and quality of the projects using statistical techniques such as descriptive statistics, histogram, and scatter plots.
  • Extracted, interpreted, and analyzed trends or patterns in complex data sets using advanced excel to create models and dashboards.
  • Hands on experience inAzureDevelopment, worked onAzure web application,App services,Azure storage,Azure SQL Database,Virtual machines.
  • Extraction of large amounts ofdata for analysis and reporting. Responsible for documentation of all analysis as well asdata discrepancies in both Spark and Python.
  • Working closes with other analysts to reconcile all issues related todata production, data extraction and delivery in order to ensure the integrity of thedata and the reporting that it is used for.
  • Determined the missing data, outlier and invalid data and applied appropriate data management techniques.
  • Worked on data manipulation and raw marketing data of different formats from multiple sources and prepared the data for Sentiment analysis of all the customer medical issue data using packages like NLP with NLTK (Natural Language Processing with Python / Analyzing Text with the Natural Language Toolkit).
  • Wrote script in python to predict number of people getting effect of some diseases, by collecting set of predicted (symptoms) data from all medical sectors and evaluated with outcome data and Make the aware of people using Machine Learning Module like logistic regression.
  • Designed custom reports, charts, tables and dashboards using AWS Quicksight, Power BI and for marketing and operations teams for their decision-making process.
  • Participates in technical design reviews to ensure user requirements are met by QA testing
  • Create focused reports and dashboards on content/channel performance, lead generation and conversion rates.
  • Assisted in writing a wide range of documents including work plans, monthly, quarterly and annual progress reports, and provider/grantee guidance materials and manual.

Environment: R studio, R programming, Jupyter Notebook,Teradata, Alteryx 11.0 & 11.3, PYTHON, AWS Quicksight, Sentiment analysis, My SQL, CRM, Tableau, MS Access, Power BI.

Confidential, TX

Data Scientist

Responsibilities:

  • Translated business challenges into math/statistical hypotheses and prepare analysis plans to highlight potential value for the business
  • Worked on gathering the requirements needed for the futher data analysis and dvelopement stages.
  • Research and development of machine learning/Deep Learning models in risk management focusing on Claim Financials, Indemnity Benefits and Return to Work areas
  • Responsible for deriving actionable insights through interpretation of model predictions and analyzing worker compensation claims data for profitable actions
  • Develop, recommend and implement procedural changes to increase the effectiveness and accuracy of production
  • Developed natural language generation models based on Aritificial Intelligence e.g. RNN and Variational AutoEncoder for generation of future notes for worker compensation claims using Keras/Tensorflow
  • Utilized Java, Eclipse/ J2EE, tested programing code & Applications, collaborated with team and management.
  • Managed and analyzed time series data and other large sets using statistical tools and techniques to answer business questions
  • Worked with the software engineering team to integrate the Machine Learning solution in production environment
  • Worked on big data using PySpark and performed machine learning using ML, MLlib packages
  • Trained machine learning/deep learning models in Google Cloud Platform and AWS
  • Implemented an AutoEncoder based anomaly detection system to find unusual claims for further investigations

Environment: MySQL, Oracle, Cassandra, Linux, SQL, Spark, Keras, TensorFlow, AWS, Windows 10, MS Excel,QlickView, Big data, VBA, R, Python, Github, Google Cloud, Kubernetes, Outlook, JIRA.

Confidential

Data Analyst

Responsibilities:

  • Performed taks like Identify, extract, clean, transform, validate and model data across multiple domains using various technologies and languages
  • Performed Data Cleaning Using R studio.
  • Utilized new and advanced methods to extract, transform and load data
  • Worked on SQL, SQL Reporting Services, ODBC, MS Access, and MS Excel and/or equivalent data capture/reporting tools
  • Used SQL skills to identify data anomalies which may impact our processes
  • Build requirements for, develop and enhance automated processes that provide functional areas with data and information
  • Analyzed the effectiveness of programs, summarize and present data to enhance future efforts using appropriate statistical methods
  • Developed business relationships and partnerships with cross-functional targeted business owners and senior leaders across the organization
  • Collaborated with cross-functional teams responsible for implementing projects in support of divisional and firm-wide business objectives
  • Utilized a strong understanding of the Edward Jones Solutions-Based Approach to investing that includes client needs, client goals, and tailored solutions

Environment: Teradata, MySQL, R Studio, MS Access, Ms Excel

Confidential

Data Analyst

Responsibilities:

  • Wrote structured application/interface code from specifications conforming to established methodology and standards. Provide first-line problem resolution or escalation, as well as technical assistance.
  • Analyze end user needs and developing customer oriented solutions which interface with existing applications. Designed and analyzed theSQLServer database and involved in gathering the user requirements.
  • Wrote simple and advanced SQL queries and scripts to create standard reports for senior managers.
  • Imported data from SQL Server and Excel to SAS datasets. Performed data research forecasting models into production by using Base SAS, SAS Macros, PROC SQL andmany other SAS tools.
  • Designed several ETL jobs using SAS DI Studio to port the ODS data into data marts with tasks involving configuring metadata libraries in SCM.
  • Improved financial status by analyzing results, monitoring variances, identifying trends, recommending actions to management with different time and place using time Series Analysis.
  • Imported the customerdata into Python using Pandas libraries and performed variousdata analysis found patterns indata which helped in key decisions.
  • Maintain data integrity across modules (i.e. marketing, Sales, Contact) in HRMS/CRM Database using Machine learning modules like Regression, Logistics and cluster Modules.
  • Provide data analysis, finding Root/Cause data integrity problems in HRMS, Report generation. Ensured all data has a complete and accurate definition
  • Implemented Python libraries such as Numpy, Matplotlib, Pandas, SKlearn and used them to create dashboards and visualizations, with using IDE - Spyder/Jupyter notebook.
  • Determined the cost of operations by establishing standard costs, collecting operationaldata and by using research analytics anddata modeling techniques.
  • Guided cost analysis process by establishing and enforcing policies and procedures, providing trends and forecasts, explaining processes and techniques, recommending actions by Regression, Logistic and cluster modules.
  • Created audit reports involving charts and graphs using Excel, Power BI dashboards, to illustrate revenue comparison for subprime lending.
  • Responsible for data aggregation, data pre-processing, missing value imputation and descriptive and inferential analysis.
  • Worked with the QA team to prepare test strategy documents which include testing overview approach, strategy, role, responsibilities and complete scope of testing.

Environment: python, Jupyter notebook, SQL, Microsoft Office, MS Power BI, MS Dynamics CRM, SQL Server Reporting Services (SSRS), Microsoft Excel, Spyder.

We'd love your feedback!