Data Scientist Resume
Minneapolis, MN
SUMMARY
- Above 8 years of experience in Machine Learning, Datamining with large datasets of Structured and Unstructured data, Data Acquisition, DataValidation, Predictive modeling, Data Visualization.
- Experience in coding SQL/PL SQL using Procedures, Triggers and Packages.
- Extensive experience in Text Analytics, developing different Statistical Machine Learning, Data Mining solutions to various business problems and generating data visualizations using R, Python.
- Excellent Knowledge in Relational Database Design, Data Warehouse/OLAP concepts and methodologies.
- Expertise in transforming business requirements into analytical models, designing algorithms, building models, developing data mining and reporting solutions that scales across massive volume of structured and unstructured data.
- Experience with Big Data technologies like Hadoop & Spark would be a plus.
- Experience working at Pricing and/or Revenue Management would be a plus.
- Experience in multiple software tools and languages to provide data - driven analytical solutions to decision makers or research teams.
- Good Knowledge in NoSQL databases like MongoDB and HBase.
- Inferential Statistics and Descriptive, Hypothesis Testing and Sampling.
- Time Series Analysis -ARIMA, Neural Networks, Sentiment Analysis, Forecasting and Text Mining.
- Develop, maintain and teach new tools and methodologies related to data science and high performance computing.
- Mapping and tracingdatafrom system to system in order to establishdatahierarchy andlineage.
- Familiarity with agile principles (e.g. Scrum), facilitating workshops and prototyping.
- Cluster Analysis, Principal Component Analysis (PCA), Association Rules, Recommender Systems.
- Hands on experience in credentials and experience in database management and data visualization.
- Strong experience in Software Development Life Cycle (SDLC) including Requirements Analysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
- Adept in statistical programming languages like R and also Python including Big Data technologies like Hadoop, Hive.
- Hands on experience with RStudio for doing data preprocessing and building machine learning algorithms on different datasets.
- Collaborated with the lead Data Architect to model the Data warehouse in accordance to FSLDM subject areas, 3NF format, Snow flake schema.
- Worked and extracteddatafrom various database sources like Oracle, SQL Server and DB2.
- Implemented machine learning algorithms on large datasets to understand hidden patterns and capture insights.
- Excellent communication skills (verbal and written) to communicate with clients and team, prepare + deliver effective presentations.
- Predictive Modeling Algorithms: Logistic Regression, Linear Regression, Decision Trees, K-Nearest Neighbors, Bootstrap Aggregation (Bagging), Naive Bayes Classifier, Random Forests, Boosting, Support Vector Machines.
TECHNICAL SKILLS
Database Design Tools and Data Modeling: Fact & Dimensions tables, physical & logical data modeling, Normalization and De-normalization techniques, Kimball.
Databases: SQL Server 20017, MS-Access, Oracle 11g, Sybase and DB2.
Languages: PL/SQL, SQL, T-SQL, C, C++, XML, HTML, DHTML, HTTP, Matlab, Python.
Tools: and Utilities: SQL Server 2016/2017, SQL Server Enterprise Manager, SQL Server Profiler, Import & Export Wizard, Visual Studio v14, .Net, Microsoft Management Console, Visual Source Safe 6.0, DTS, Crystal Reports, Power Pivot, ProClarity, Microsoft Office 2007/10/13, Excel Power Pivot, Excel Data Explorer, Tableau 8/10, JIRA
Operating Systems: Microsoft Windows 8/7/XP, Linux and UNIX
PROFESSIONAL EXPERIENCE
Confidential, Minneapolis MN
Data Scientist
Responsibilities:
- Transformation of data using SSIS.
- Build analytical solutions and models by manipulating large data sets.
- Implementing data mining and statistical machine learning solutions to various business problems.
- Applied machine learning and statistical techniques to large datasets to find actionable insights.
- Involved in complete Software Development Life Cycle (SDLC) process by analyzing business requirements and understanding the functional work flow of information from source systems to destination systems.
- Played critical role in collecting data from different data sources and data system like SAP, JDE, Lubes, Hadoop, etc.
- Created ETL packages using SSIS to extract data from relational database and then transform and load into the data mart.
- Transforming and merging all the weekly client data into yearly file using ETL SSIS
- Used Visual Team Foundation server for version control, source control and reporting.
- KT with the client to understand their various Data Management systems and understanding the data.
- Creating meta-data and data dictionary for the future data use/ data refresh of the same client.
- Structuring the Data Marts to store and organize the customer’s data.
- Mapping flow of trade cycle data from source to target and documenting the same.
- Performing QA on the data extracted, transformed and exported to excel.
Environment: Hadoop, Oracle 11g, MS Office, SSMS, SSIS, Power BI, Microsoft reporting tools, Big Data.
Confidential, MI
Data Scientist
Responsibilities:
- A highly immersive Data Science program involving Data Manipulation & Visualization, Web Scraping, Machine Learning, Python programming, SQL, GIT, Unix Commands, NoSQL, MongoDB, Hadoop.
- Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
- Used pandas, numpy, seaborn, scipy, matplotlib, scikit-learn, NLTK in Python for developing various machine learning algorithms.
- Installed and used Caffe Deep Learning Framework
- Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python.
- Worked asDataArchitectsand ITArchitectsto understand the movement ofdataand its storage and ER Studio 9.7
- Participated in all phases of data mining; data collection, data cleaning, developing models, validation, visualization and performed Gap analysis.
- Data Manipulation and Aggregation from different source using Nexus, Toad, Business Objects, Power BI and Smart View.
- Implemented Agile Methodology for building an internal application.
- Focus on integration overlap and Informatica newer commitment to MDM with the acquisition of Identity Systems.
- Good knowledge of Hadoop Architecture and various components such as HDFS, JobTracker, Task Tracker, NameNode, DataNode, Secondary NameNode, and MapReduce concepts.
- AsArchitectdelivered various complex OLAP databases/cubes, scorecards, dashboards and reports.
- Programmed a utility in Python that used multiple packages (scipy, numpy, pandas)
- Implemented Classification using supervised algorithms like Logistic Regression, Decision trees, KNN, Naive Bayes.
- Used Teradata15 utilities such as Fast Export, MLOAD for handling various tasks data migration/ETL from OLTP Source Systems to OLAP Target Systems
- Experience in Hadoop ecosystem components like Hadoop MapReduce, HDFS, HBase, Oozie, Hive, Sqoop, Pig, Flume including their installation and configuration.
- Updated Python scripts to match trainingdatawith our database stored in AWS Cloud Search, so that we would be able to assigneach document a response label for further classification.
Environment: regression, logistic regression, Hadoop, Teradata, OLTP, Unix, Python, MLLib, SAS, random forest, OLAP, HDFS, NLTK, SVM, JSON and XML.
Confidential, Kansas, MO
Data Scientist/Data Analyst
Responsibilities:
- Generated cost-benefit analysis to quantify the model implementation comparing with the former situation
- Worked on model selection based on confusion matrices, minimized the Type II error
- Worked on data cleaning and reshaping, generated segmented subsets using Numpy and Pandas in Python
- Wrote and optimized complex SQL queries involving multiple joins and advanced analytical functions to perform data extraction and merging from large volumes of historical data stored in Oracle 11g, validating the ETL processed data in target database
- Applied various machine learning algorithms and statistical modeling like decision tree, logistic regression, Gradient Boosting Machine to build predictive model using scikit-learn package in Python
- Developed Python scripts to automate data sampling process. Ensured the data integrity by checking for completeness, duplication, accuracy, and consistency
- Identified the variables that significantly affect the target
- Continuously collected business requirements during the whole project life cycle.
- Conducted model optimization and comparison using stepwise function based on AIC value
- Generated data analysis reports using Matplotlib, Tableau, successfully delivered and presented the results for C-level decision makers
Environment: Matplotlib, Scikit-Learn, Tableau 7, Python 2.6.8, Numpy, Pandas, MongoDB, Oracle 10g, SQL
Confidential, Richmond VA
Data Architect/Data Modeler
Responsibilities:
- Develop Integrations jobs to transfer data from source system to Hadoop.
- Installation of Talend Studio.
- Technical design documents for Transformation processes.
- Application of business rules on the data being transferred.
- Task allocation for the ETL and Reporting team.
- Communicate effectively with client and their internal development team to deliver product functionality requirements.
- Architecting and design of data warehouse ETL processes.
- Demo of POC built for the prospective customer and provide guidance and gather the feedback to backend ETL testing on SQL Server 2008 using SSIS.
- Create the Operational manual Document.
- Create Integration Jobs to backup a copy of data in network file system.
- Design and implement the ETL Data model and create staging, source and Target tables in SQL server database.
- Gathering and analysis requirements definition meetings with business users and document meeting outcomes.
Environment: ETL, ODS, Hadoop, MS Office, Talend Studio, OLAP, SQL Server 2008.
Confidential
Data Analyst/Data Modeler
Responsibilities:
- Worked with data compliance teams, Data governance team to maintain data models, Metadata, Data Dictionaries; define source fields and its definitions.
- Participated in JAD session with business users and sponsors to understand and document the business requirements in alignment to the financial goals of the company.
- Involved in analysis of Business requirement, Design and Development of High level and Low level designs, Unit and Integration testing
- Performed second and third normalizations for ER data model of OLTP system
- Designed, Build the Dimensions, cubes with star schema and Snow Flake Schema using SQL Server Analysis Services (SSAS).
- Design and model the reporting data warehouse considering current and future reporting requirement
- Performed data analysis and data profiling using complex SQL on various sources systems including Teradata, SQL Server.
- Developed the logical data models and physical data models that confine existing condition/potential status data fundamentals and data flows using ER Studio
- Created the conceptual model for the data warehouse using Erwin data modeling tool.
- Reviewed and implemented the naming standards for the entities, attributes, alternate keys, and primary keys for the logical model.
- Translate business and data requirements into Logical data models in support of Enterprise Data Models, ODS, OLAP, OLTP, Operational Data Structures and Analytical systems.
Environment: Oracle, ODS, OLAP, SQL*Loader, PL/SQL, Teradata, SQL Server 2008, OLTP, Informatica Power Center.
Confidential
Data Analyst
Responsibilities:
- Generated periodic reports based on the statistical analysis of the data using SQL Server Reporting Services.
- Developed and executed load scripts using Teradata client utilities MULTILOAD, FASTLOAD and BTEQ.
- Responsible for development and testing of conversion programs for importing Data from text files into map Oracle Database utilizing PERL shell scripts & SQL*Loader.
- Used Graphical Entity-Relationship Diagramming to create new database design via easy to use, graphical interface.
- Applied Business Objects best practices during development with a strong focus on reusability and better performance.
- Developed Tableau visualizations and dashboards using Tableau Desktop.
- Designed different type of STAR schemas for detailed data marts and plan data marts in the OLAP environment.
- Maintained metadata (data definitions of table structures) and version controlling for the data model.
- Developed SQLscripts for creating tables, Sequences, Triggers, views and materialized views.
- Utilized Erwin's forward/reverse engineering tools and target database schema conversion process.
- Worked on creating enterprise wide Model EDM for products and services in Teradata Environment based on the data from PDM. Conceived, designed, developed and implemented this model from the scratch.
- Co-ordinate with various business users, stakeholders and SME to get Functional expertise, design and business test scenarios review, UAT participation and validation of financial data.
- Write SQL scripts to test the mappings and Developed Traceability Matrix of Business Requirements mapped to Test Scripts to ensure any Change Control in requirements leads to test case update.
Environment: MS SQL Server, Oracle SQL Developer, PL/SQL, Business Objects, Windows XP, TOAD, SQL*PLUS, SQL*LOADER, Tableau, Business Objects, Informatica, XML.
