Data Scientist Resume
Boston, MA
SUMMARY:
- 7+ years of experience in strong engineering, management and programming background, especially in mathematics, data analysis, programming and design, build, deploying machine learning applications to solve real - world problems empirically.
- Skilled in Machine learning, Statistics, management, problem-solving, and programming.
- Adept in statistical programming languages like Python, R and SQL.
- Excellent analytical, data mining, data collecting, data processing and cleaning, feature extraction, data visualization skills.
- Excellent ability to understand business data needs and apply complex technology in data analyzing and modeling to create data-driven decision support systems and to explain them to business people.
- Successfully interpreted data to draw conclusions for managerial action and strategy.
- Done course work related to computer science: Data Structures, Application of Neural Network, Data Base Management Systems.
- Done MOOCs on Machine Learning and Data Analysis in coursera taught by Andrew NG, Stanford University.
- Great knowledge in Data Analysis, Risk Management, Lean Manufacturing, Six Sigma, Supply Chain Management, Planning & Forecasting, Inventory Management and Process Improvement principles.
- Developed statistical models in R and Python using various supervised and unsupervised machine learning algorithms such as Linear Regression and Logistic Regression, Classification, Decision Trees, Gradient Boosting, KNN, Support Vector Machines, Naive Bayes, K-Means Clustering, Neural Networks, Principal Component Analysis and Recommender Systems on structured and unstructured data.
- Good domain knowledge in Insurance, Finance, Health care etc.
- Created statistical models for the collected data, exploratory, pre-processing, to provide conclusions which guides for the policy making decisions.
- Experience in writing Sub Queries, Stored Procedures, Triggers, Cursors, and Functions in MySQL.
- Experienced in Agile methodology and SCRUM process.
- Hands on experience in design, management and visualization of databases using Oracle, MySQL, and SQL Server.
- By using Lean Six Sigma tools and machine learning algorithms improved standardized claims submission, reduced claims escalations and improved customer satisfaction.
- Experienced in using a combination of data visualization tools, Tableau and communication tools to clearly and effectively explain the problem, root cause and recommendations.
- Done various projects like loan prediction, house price prediction, speech recognition, text classification and hand-written digit recognition by using R, Python, NLP, data analysis, management, neural networks and programming skills.
- Successfully interpreted data to draw conclusions for managerial action and strategy.
- Good industry knowledge, analytical & problem-solving skills and ability to work well with in a team as well as an individual.
- Highly creative, innovative, committed, intellectually curious, business savvy with good communication and interpersonal skills.
- Ability and Adaptability to work with different technologies on different platforms.
PROFESSIONAL EXPERIENCE:
Confidential, Boston, MA
Data Scientist
Responsibilities:
- Worked as a Data Modeler/Analyst to generate Data Models using Erwin and developed relational database system. A highly immersive Data Science program involving Data Manipulation & Visualization, Web Scraping, Machine Learning, SQL, GIT, Unix Commands, NoSQL, MongoDB.
- Provide expertise and recommendations for physical database design, architecture, testing, performance tuning and implementation.
- Transformed Logical Data Model to Erwin, Physical Data Model ensuring the Primary Key and Foreign Key relationships in PDM, Consistency of definitions of Data Attributes and Primary Index Considerations.
- Extensively worked on Data Modeling tools Erwin Data Modeler to design the data models.
- Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
- Data was visualized using different visualization (scatter plot, box plots, and histograms) techniques from ggplot2 package in R.
- Designed the physical model for implementing the model into theoracle9i physical database.
- Involved with Data Analysis Primarily Identifying Datasets, Source Data, Source Meta Data, Data Definitions and Data Formats
- Application of various machine learning algorithms and statistical modelling like Logistic Regression, Decision tree, SVM, to identify Volume using scikit-learn package in R & Python.
- Improve efficiency and accuracy by evaluating model in R.
- Designed tables and implemented the naming conventions for Logical and Physical Data Models in Erwin 7.0. Used R to manipulate data, develop and validate quantitative models.
- Cleansed the data by eliminating duplicate and inaccurate data in R and Python.
- Implementing Neural network for Large-scale adoption of Electronic Health Records (EHR) using H2O, Tensor Flow, OpenCV, and Keras.
- Solution architecting BIG Data solution for Projects & Proposal using Hadoop, Spark, ELK Stack, Kafka, Tensor flow.
- Built text analytics solutions with IBM Watson and Google Cloud which in turn saved analysts operational and processing time.
- Designed logical and physical data models for multiple OLTP and Analytic applications.
- Extensively used the Erwin design tool &Erwin model manager to create and maintain the Data Mart.
- Performance tuning of the database, which includes indexes, and optimizing SQL statements, monitoring the server.
- Utilized Power BI (Power View) to create various analytical dashboards that depicts critical KPIs such as legal case matter, billing hours and case proceedings along with slicers and dicers enabling end-user to make filters
- Utilized Power Query in Power BI to Pivot and Un-pivot the data model for data cleansing and data massaging.
- Wrote simple and advanced SQL queries and scripts to create standard and Adhoc reports for senior managers.
- Collaborated the data mapping document from a source to target and the data quality assessments for the source data.
- Created S3 buckets and managed roles and policies for S3 buckets. Utilized S3 buckets and Glacier for file storage and backup on AWS cloud. Used Dynamo DB to store the data for metrics and backend reports.
- Worked with Elastic Beanstalk for quick deployment of services such as EC2 instances, Load balancer, and databases on the RDS on the AWS cloud environment.
- Used Java code to connect AWS S3 buckets by using AWS SDK, to access media files related to the application.
- Used Amazon Simple Workflow service (SWF) for data migration in data centers which automates the process and tracks every step and logs are maintained in S3 bucket.
- Performed Data Analysis and Data Profiling and worked on data transformations and data quality rules.
- Created SSIS Packages using Pivot Transformation, Execute SQL Task, Data Flow Task, etc. to import data into the data warehouse.
- Developed and implemented SSIS, SSRS and SSAS application solutions for various business units across the organization.
Environment: SQL, GIT, Unix Commands, IBM Watson, Tensor flow, NoSQL, R/R studio, Python, Tableau, Mongo DB, SSIS, SSRS,SSAS, AWS,S3,EC2,RDS,SWF,Dynamo DB, Glacier, Erwin, Tableau, OBIEE.
Confidential - Irving, TX
Data Scientist
Responsibilities:
- Used Tableau to automatically generate reports. Worked with partially adjudicated insurance flat files, internal records, 3rd party data sources, JSON, XML and more.
- Experienced in building models by using Spark (PySpark, SparkSQL, Spark MLLib, and Spark ML).
- Experienced in Cloud Services such as AWS EC2, EMR, RDS, S3 to assist with big data tools, solve the data storage issue and work on deployment solution.
- Worked with several R packages including knitr, dplyr, SparkR, Causal Infer, spacetime.
- Performed Exploratory Data Analysis and Data Visualizations using R, and Tableau.
- Implemented end-to-end systems for Data Analytics, Data Automation and integrated with custom visualization tools using R, Mahout, Hadoop and MongoDB.
- Gathering all the data that is required from multiple data sources and creating datasets that will be used in analysis.
- Knowledge extraction from Notes using NLP (Python, NLTK, MLLIB, PySpark,)
- Independently coded new programs and designed Tables to load and test the program effectively for the given POC's using with Big Data/Hadoop.
- Worked with BTEQ to submit SQL statements, import and export data, and generate reports in Teradata.
- Built and optimized data mining pipelines of NLP, and text analytic to extract information.
- Coded R functions to interface with Caffe Deep Learning Framework
- Working in Amazon Web Services cloud computing environment
- Interacted with the other departments to understand and identify data needs and requirements and work with other members of the IT organization to deliver data visualization and reporting solutions to address those needs.
- Perform a proper EDA, Univariate and bi-variate analysis to understand the intrinsic effect/combined effects.
- Designed data models and data flow diagrams using Erwin and MS Visio.
- Established Data architecture strategy, best practices, standards, and roadmaps.
- Performed data cleaning and imputation of missing values using R.
- Developed, Implemented & Maintained the Conceptual, Logical & Physical Data Models using Erwin for Forward/Reverse Engineered Databases.
- Built and optimized data mining pipelines of NLP, and text analytic to extract information.
- Worked with Hadoop eco system covering HDFS, HBase, YARN and Map Reduce.
- Creating customized business reports and sharing insights to the management.
- Take up ad-hoc requests based on different departments and locations.
- Used Hive to store the data and perform data cleaning steps for huge datasets.
- Created dash boards and visualization on regular basis using ggplot2 and Tableau.
Environment: R 3.0, Erwin 9.5, Tableau 8.0, MDM, Qlikview, MLLib, PL/SQL, HDFS, Teradata 14.1, JSON, HADOOP (HDFS), Map Reduce, PIG, Spark, R Studio, MAHOUT, JAVA, HIVE, AWS.
Confidential, Atlanta, GA
Data Analyst
Responsibilities:
- Worked for Risk management team in identifying the risk involved in the Mortgage process by evaluating the customer and property records.
- Involved in addressing a wide range of challenging problems using techniques from applied statistics, machine learning and data mining fields.
- Involved in providing insights using Machine learning algorithms by tailoring to particular needs and evaluated on large data sets using Caret, e1071, rpart, randomForest, glmnet, gbm, mboost and arules in R.
- Integrated new tools and developed technology frameworks/prototypes to accelerate the data integration process and empower the deployment of predictive analytics by developing Spark Scala modules with R.
- Experience in descriptive statistics and hypothesis testing using Chi-square, T-test, Pearson correlation and Analysis of variance (ANOVA).
- Advanced knowledge of statistical techniques in Sampling, Probability, Multivariate data analysis, PCA, and Time-series analysis using SAS.
- Experience in transferring and managing data to Hadoop clusters using Kafka, Sqoop, Oozie and Zookeeper.
- Involved in fixing invalid mappings, testing of Stored Procedures and Functions, Unit and Integrating testing of Informatica Sessions, Batches, and the Target Data.
- Responsible for creating Hive tables, loading the structured data resulted from Map Reduce jobs into the tables and writing hive queries to further analyze the logs to identify issues and behavioral patterns.
- Analyzed Relational & Non-relational data using MySQL and HBase.
- Contributed Technology, Project management, and Business management functions to push the business forward with innovative solutions.
- Performed data acquisition and exploratory data analysis in R.
- Visualized team metrics and communicated to the higher management using PowerBI.
Environment: R, MySQL, Hive, HBase, Spark, PowerBI and GIT.
Confidential - Bellevue, WA
Data Analyst
Responsibilities:
- Performed Data Profiling to learn about behavior with various features such as traffic pattern, location, and time, Date and Time etc.
- Created ecosystem models (e.g. conceptual, logical, physical, canonical) that are required for supporting services within the enterprise data architecture (conceptual data model for defining the major subject areas used, ecosystem logical model for defining standard business meaning for entities and fields, and an ecosystem canonical model for defining the standard messages and formats to be used in data integration services throughout the ecosystem).
- Used Pandas, NumPy, seaborn, SciPy, Matplotlib, Scikit-learn, NLTK in Python for developing various machine learning algorithms and utilized machine learning algorithms such as linear regression, multivariate regression, naive Bayes, Random Forests, K-means, &KNN for data analysis.
- Integrated new tools and developed technology frameworks/prototypes to accelerate the data integration process and empower the deployment of predictive analytics by developing Spark Scala modules with R.
- Wrote several Teradata SQL Queries using Teradata SQL Assistant for Ad Hoc Data Pull request.
- Responsible for business case analysis, requirements gathering use case documentation, prioritization and product/portfolio strategic roadmap planning, high level design and data model.
- Oversee development of all Tableau dashboards for organization.
- Coordinate data delivery from other developers, in order to update dashboards on a monthly basis.
- Experience parsing data stored in Excel, CSV, JSON, HTML, PDF, TXT, and other file formats.
- Finished project focusing on predicting blood born infections in patients, after undergoing surgery.
- Built object-oriented framework to easily allow construction of multi-layer ensemble machine-learning models, using Scikit-learn, XG Boost, Theano, and other Python toolkits.
- Developed Java application to extract text features from hundreds of thousands of clinical encounters.
- Developed simulations to study and understand the effects of CMS bundled payment model on hospital output using claims data and hospital accounting data.
- Learned HTML, CSS, and JavaScript to develop a web application to demonstrate organizations.
- Built website to act as a code repository for all the organizations parsers
Environment: R 3.0, Erwin 9.5, Tableau 8.0, MDM, QlikView, MLLib, PL/SQL, HDFS
Confidential
Data Analyst
Responsibilities:
- Worked with business requirements analysts/subject matter experts to identify and understand requirements. Conducted user interviews and data analysis review meetings.
- Defined key facts and dimensions necessary to support the business requirements along with Data Modeler.
- Created draft data models for understanding and to help Data Modeler.
- Resolved the data related issues such as: assessing data quality, data consolidation, evaluating existing data sources.
- Manipulating, cleansing & processing data using Excel, Access and SQL.
- Responsible for loading, extracting and validation of client data.
- Coordinated with the front end design team to provide them with the necessary stored procedures and packages and the necessary insight into the data.
- Participated in requirements definition, analysis and the design of logical and physical data models
- Leading data discovery discussions with Business in JAD sessions and map the business requirements to logical and physical modeling solutions.
- Conducted data model reviews with project team members captured technical metadata through data modeling tools
- Code standard Informatica ETL routines developed standard Cognos Reports.
- Collaborated with ETL teams to create data landing and staging structures as well as source to target mapping documents
- Ensure data warehouse database designs efficiently support BI and end user requirements.
- Collaborated with application and services teams to design databases and interfaces which fully meet business and technical requirements
- Maintain expertise and proficiency in the various application areas.
- Maintain current knowledge of industry trends and standards.
Environment: Informatica 9.1, Oracle 11g, SQL Developer, PL/SQL, Cognos, Splunk, TOAD, MS Access, MS Exce
