Data Scientist Resume
TN
SUMMARY
- Over 8+ years of experience in Machine Learning, Data mining with large datasets of Structured and Unstructured data, Data Acquisition, Data Validation, Predictive modelling, Data Visualization.
- Excellent Knowledge in Relational Database Design, Data Warehouse/OLAP concepts and methodologies.
- Collaborated with the lead Data Architect to model the Data warehouse in accordance to FSLDM subject areas, 3NF format, Snow flake schema.
- Worked and extracteddatafrom various database sources like Oracle, SQL Server and DB2.
- Mapping and tracingdatafrom system to system in order to establishdatahierarchy andlineage.
- Experience in coding SQL/PL SQL using Procedures, Triggers and Packages.
- Extensive experience in Text Analytics, developing different StatisticalMachineLearning, DataMining solutions to various business problems and generating datavisualizations using R, Python.
- Expertise in transforming business requirements into analyticalmodels, designingalgorithms, buildingmodels, developing datamining and reportingsolutions that scales across massive volume of structured and unstructured data.
- Experience with Big Data technologies like Hadoop & Spark would be a plus.
- Expertise in developing advanced statistical and econometric models to predict, quantify or forecast various operational and performance metrics.
- Experience working at Pricing and/or Revenue Management would be a plus.
- Familiarity with agile principles (e.g. Scrum), facilitating workshops and prototyping.
- Excellent communication skills (verbal and written) to communicate with clients and team, prepare + deliver effective presentations.
- Strong experience in Software Development Life Cycle (SDLC) including RequirementsAnalysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
- Adept in statisticalprogramminglanguages like R and also Python including BigData technologies like Hadoop, Hive.
- Highly skilled in using visualization tools like Tableau for creating dashboards.
- Implemented Optimization techniques for better performance on the ETL side and also on the database side.
- Experience in using Statistical procedures and Machine Learning algorithms such as ANOVA, Clustering and Regression and Time Series Analysis to analyze data for further Model Building.
- Designing of Physical Data Architecture of New system engines.
- Expertise in transforming business requirements into analytical models, designing algorithms, building models, developing datamining and reporting solutions that scales across massive volume of structured and unstructured data
- Experience in applying PredictiveModeling and MachineLearning algorithms for Analytical projects.
- Developing Logical Data Architecture with adherence to Enterprise Architecture.
- Experience in handling managing and analyzing large data sets with the team for Analytics Projects to identify key hidden insights within the Data so as to make important recommendations to the teams to enhance Objectives.
- Knowledge of working with Proof of Concepts (PoC’s) and gap analysis and gathered necessary data for analysis from different sources, prepared data for data exploration using data munging and Teradata.
- Designing of Physical Data Architecture of New system engines
- Experience and Technical proficiency in Designing,DataModeling Online Applications, Solution Lead for ArchitectingDataWarehouse/Business Intelligence Applications.
- Well experienced in Normalization and De - Normalization techniques for optimum performance in relational and dimensional database environments.
TECHNICAL SKILLS
Languages: PL/SQL, SQL, T-SQL, C, C++, XML, HTML, DHTML, HTTP, Mat lab, Python
Databases: SQL Server 2017/2016/2014, MS-AccessOracle 11g/10g/9i, Sybase 15.02 and DB2 2016.
Big Data Tools: Hadoop 2.7.2, Hive, Spark2.1.1, Pig, HBase, Sqoop, Flume
DWH / BI Tools: Microsoft Power BI, SSIS, SSRS, SSAS, Business Intelligence Development Studio (BIDS), Visual Studio, SAP Business Objects, SAP SE v 14.1(Crystal Reports) and Informatica 6.1.
Database Design Tools and Data Modeling: Star Schema/Snowflake Schema modeling, Fact & Dimensions tables, physical & logical data modeling, Normalization and De-normalization techniques, Kimball & Inmon Methodologies
Tools: and Utilities: SQL Server 2016/2017, SQL Server Enterprise Manager, SQL Server Profiler, Import & Export Wizard, Visual Studio v14, .Net, Microsoft Management Console, Visual Source Safe 6.0, DTS, Crystal Reports, Power Pivot, ProClarity, Microsoft Office 2007/10/13, Excel Power Pivot, Excel Data Explorer, Tableau 8/10, JIRA
DataModeling Tools: Erwin r9.6/9.5, ER/Studio 9.7, Star-Schema Modeling, Snowflake-Schema Modeling, FACT and dimension tables, Pivot Tables
Operating Systems: Microsoft Windows 8/7/XP, Linux and UNIX
PROFESSIONAL EXPERIENCE
Confidential, TN
Data Scientist
Responsibilities:
- As an Architect design conceptual, logical and physical models using Erwin and build datamarts using hybrid Inmon and Kimball DW methodologies.
- Assisted the project with Python programming, coding and running QA on the same from time to time.
- Defined accountability procedures governing data access, processing and storage, retention, reporting and auditing measuring contract compliance.
- Ensured that Business, Data Governance and Integration team leads are deeply involved in critical design issues and decisions.
- Created the Enterprise Information Group encompassing Business Intelligence Center of Excellence, Data Governance Working Group, Enterprise Data Warehousing and Master Data Management.
- Worked with several R packages including knitr, dplyr, SparkR, CausalInfer, spacetime.
- Implemented end-to-end systems for DataAnalytics, DataAutomation and integrated with custom visualization tools using R,Mahout, Hadoop and MongoDB.
- Gathering all the data that is required from multiple data sources and creating datasets that will be used in analysis.
- Worked withDatagovernance,Dataquality,datalineage,Dataarchitectto design various models and processes.
- Independently coded new programs and designed Tables to load and test the program effectively for the given POC's using with BigData/Hadoop.
- Designeddatamodels anddataflow diagrams using Erwin and MS Visio.
- As an Architect implemented MDM hub to provide clean, consistent data for a SOA implementation.
- Developed, Implemented & Maintained the Conceptual, Logical & PhysicalDataModels using Erwin for Forward/Reverse Engineered Databases.
- EstablishedDataarchitecture strategy, best practices, standards, and roadmaps.
- Lead the development and presentation of a data analytics data-hub prototype with the help of the other members of the emerging solutions team
- Performed data cleaning and imputation of missing values using R.
- Worked with Hadoop eco system covering HDFS, HBase, YARN and Map Reduce.
- Involved in business process modeling using UML
Environment: Teradata 13.1, Informatica 6.2.1, Ab Initio, Business Objects, Oracle 9i, PL/SQL, Microsoft Office Suite (Excel, Vlookup, Pivot, Access, Power Point), Visio, VBA, Micro Strategy, Tableau, ERWIN.
Confidential, San Jose, CA
Data Scientist
Responsibilities:
- Responsible for performing Machine-learning techniques regression/classification to predict the outcomes.
- Responsible for design and development of advanced R/Python programs to prepare transform and harmonize data sets in preparation for modelling.
- Identifying and executing process improvements, hands-on in various technologies such as Oracle, Informatica, Business Objects.
- Designed the prototype of theDatamart and documented possible outcome from it for end-user.
- Involved in business process modelling using UML
- Developed and maintaineddatadictionary to create metadata reports for technical and business purpose.
- Handled importing data from various data sources, performed transformations using Hive, Map Reduce, and loaded data into HDFS.
- Interaction with Business Analyst, SMEs and other Data Architects to understand Business needs and functionality for various project solutions.
- Created SQL tables with referential integrity and developed queries using SQL, SQL*PLUS and PL/SQL.
- Worked closely with business, datagovernance, SMEs and vendors to define data requirements.
- Design, coding, unit testing of ETL package source marts and subject marts using Informatica ETL processes for Oracle database.
- Identifying and executing process improvements, hands-on in various technologies such as Oracle, Informatica, and Business Objects.
- Designed both 3NF data models for ODS, OLTP systems and dimensional data models using Star and Snow Flake Schemas.
- Developed large data sets from structured and unstructured data. Perform data mining.
- Partnered with modellers to develop data frame requirements for projects.
- Researched, evaluated, architected, and deployed new tools, frameworks, and patterns to build sustainable Big Data platforms for the clients.
- Performed Ad-hoc reporting/customer profiling, segmentation using R/Python.
- Involved in the daily maintenance of the database that involved monitoring the daily run of the scripts as well as troubleshooting in the event of any errors in the entire process.
Environment: Erwin 8, Teradata 13, SQL Server 2008, Oracle 9i, SQL*Loader, PL/SQL, ODS, OLAP, OLTP, SSAS, Informatica Power Center 8.1.
Confidential, Wisconsin
Data Scientist
Responsibilities:
- Created statistical models for the collected data, exploratory, pre-processing, in order to provide conclusions which guides for the policy making decisions.
- Worked with DBMI, SAS IT, BI team and automated 2 mature modeling and scoring process.
- Monitored study results and interpret data in a manner which provides clarity and an undeniable conclusion.
- Used SAS CI studio and SAS MO for the optimization of the campaign performance.
- Changed our model codes and results for smooth connection with SAS MO system and managed the testing, implementation of MO.
- Developed automation process for variable and model selection on over 5,000 core SKU's.
- Communicated detailed statistical and scientific findings to lay individuals.
- Worked on Noise Reduction methods exponential smoothing and Fast Fourier Transformation methods and made comparison with regression methods.
- Carried out Regression Analysis with R/SAS, investigated on the model for problems like goodness of fit, over-fitting, Multi co linearity, residual normality, etc. Established our linear model for forecasting.
- Performed regular research and gathered new statistical evidence at every opportunity.
- Published manuscripts in scientific journals and papers presented in various conferences.
- Data analyzed for scientists, research scholars, medical doctors, and also to post.
- Utilized SPSS and Minitabsoftware to randomize, analyze and interpretation of data.
- Summarized findings, created reports, & maintained database of statistical information.
- Used advanced excel to create pivot tables, charts, used VLOOKUP, and other functions.
- EDA analysis like testing of hypotheses, generating graphs, cleaning of data, removing putlier points, transforming data, melting of data, import and export of data.
- Word Cloud Algorithm and Naive bayes applied word cloud algorithm on customer review data and found sentimental analysis, converted unstructured data to structured data, predicted insights of what and how customer want from our software.
- Published manuscripts in scientific journals and papers presented in various conferences.
- Delivered solutions using clustering algorithm, logistic regression, composite index analytics, decision tree and other algorithms as per requirement.
- Evaluation of models for Bias and variance problem. Evaluating models for best parameters using K folds.
- Achieved good prediction for different promotion media on various products.
Environment: SSMS 17.x, SSIS 2016, Power BI 2.37, Oracle 12c/11g, MS Office 2013, Microsoft reporting tools, Big Data, Hadoop 2.8.1
Confidential, Michigan
Data Scientist
Responsibilities:
- Worked as aDataModeler/Analyst to generateDataModels using Erwin and developed relational database system.
- Analyzed the business requirements of the project by studying the Business Requirement Specification document.
- Extensively worked on Data Modeling tools Erwin Data Modeler to design the data models.
- Designed mapping to process the incremental changes that exists in the source table. Whenever source data elements were missing in source tables, these were modified/added in consistency with third normal form based OLTP source database.
- Designed tables and implemented the naming conventions for Logical and Physical Data Models in Erwin 7.0.
- Provide expertise and recommendations for physical database design, architecture, testing, performance tuning and implementation.
- Designed logical and physical data models for multiple OLTP and Analytic applications.
- Extensively used the Erwin design tool & Erwin model manager to create and maintain the Data Mart.
- Involved with Data Analysis Primarily Identifying Data Sets, Source Data, Source Meta Data, Data Definitions and Data Formats
- Performance tuning of the database, which includes indexes, and optimizing SQL statements, monitoring the server.
- Wrote simple and advanced SQL queries and scripts to create standard and ad hoc reports for senior managers.
- Collaborated thedatamapping document from source to target and thedataquality assessments for the sourcedata.
- Used Expert level understanding of different databases in combinations for Data extraction and loading, joining data extracted from different databases and loading to a specific database.
Environment: Erwin r7.0, SQL Server 2000/2005, Windows XP/NT/2000, Oracle 8i/9i, MS-DTS, UML, UAT, SQL Loader, OOD, OLTP, PL/SQL, MS Visio, Informatica
Confidential
Data Architect/Data Modeler
Responsibilities:
- Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python.
- Worked asDataArchitectsand ITArchitectsto understand the movement ofdataand its storage and ERStudio9.7
- Participated in all phases of datamining; datacollection, datacleaning, developingmodels, validation, visualization and performed Gapanalysis.
- DataManipulation and Aggregation from different source using Nexus, Toad, BusinessObjects, PowerBI and SmartView.
- Implemented Agile Methodology for building an internal application.
- Focus on integration overlap and Informatica newer commitment to MDM with the acquisition of Identity Systems.
- Good knowledge of HadoopArchitecture and various components such as HDFS, JobTracker, TaskTracker, NameNode, DataNode, SecondaryNameNode, and MapReduce concepts.
- AsArchitectdelivered various complex OLAPdatabases/cubes, scorecards, dashboards and reports.
- Programmed a utility in Python that used multiple packages (scipy, numpy, pandas)
- Implemented Classification using supervised algorithms like Logistic Regression, Decision trees, KNN, Naive Bayes.
- Used Teradata15 utilities such as Fast Export, MLOAD for handling various tasks data migration/ETL from OLTP Source Systems to OLAP Target Systems
- Updated Python scripts to match trainingdatawith our database stored in AWSCloud Search, so that we would be able to assigneach document a response label for further classification.
- Datatransformation from various resources,dataorganization, features extraction from raw and stored.
- Identifying the Customer and account attributes required for MDM implementation from disparate sources and preparing detailed documentation.
- Validated the machine learning classifiers using ROC Curves and Lift Charts.
- Extracted data from HDFS and prepared data for exploratory analysis using data munging.
Environment: R 3.0, Erwin 9.5, Tableau 8.0, MDM, QlikView, MLLib, PL/SQL, HDFS, Teradata 14.1, JSON, HADOOP (HDFS), Map Reduce, PIG, Spark, R Studio, MAHOUT, JAVA, HIVE, AWS.
Confidential
Data Analyst
Responsibilities:
- Involved in Design, Development and Support phases of Software Development Life Cycle (SDLC)
- Extracted data from one or more source files for analysis.
- Worked closely with cross functional teams to encourage statistical best practices with respect to experimental design and data analysis.
- Developed product analytics, including KPIs and testing. Conducted segmentation analyses using supervised learning, clustering, and NLP.
- Accomplished multiple tasks from collecting data to organizing data and interpreting statistical information.
- Develop time series algorithms on customer transaction data, classification models on transaction types, anomaly detection, clustering using spectral analyses.
- Build machine learning pipeline for customer churn with end-to-end ETL processes with SQL/NoSQL databases.
- Lead data discovery, handling structured and unstructured data, cleaning and performing descriptive analyses, and storing as normalized tables for dashboards.
- Unearthed the raw data by doing the Exploratory Data Analysis (Classification, splitting, cross-validation).
- Used common data science toolkits, such as R, Python, NumPy, Keras, Theano, Tensorflow, etc.
- Converted raw data to processed data by merging, finding outliers, errors, trends, missing values and distributions in the data.
- Utilized various techniques like Histogram, bar plot, Pie-Chart, Scatter plot; Box plots to determine the condition of the data.
- Used Logistic Regression to obtain the probabilities for non-defaulters and defaulters.
- Identified Key performance indicators (KPI's) among all the given attributes.
- Conducted data exploration (DPLYR, GGPLOT2, Seaborn) to look for trends, patterns, grouping, and deviations in the data to understand the data diagnostics.
- Imported and exported data with Sqoop. Extracted the data to MySQL database.
- Build a Flume Agent to inject data from Directory spool source to HDFS.
Environment: Erwin r9.0, Informatica 9.0, ODS, OLTP, Oracle 10g, Hive, OLAP, DB2, Metadata, MS Excel, Mainframes MS Visio, Rational Rose, Requisite Pro, Hadoop, PL/SQL, etc.
