Data Scientist/machine Learning Engineer Resume
Seattle, WA
PROFESSIONAL SUMMARY:
- 8+ Years of e experience and comprehensive industry knowledge of Machine Learning, Statistical Modeling, Data Analytics, DataModeling, Data Architecture, Data Analysis, Data Mining, Text Mining & Natural Language Processing (NLP), Artificial Intelligence algorithms, Business Intelligence, Analytics Models (like Decision Trees, Linear & Logistic Regression,, R, Python, Spark, Scala, MS Excel, SQL and Postgre SQL, Erwin.
- Excellent knowledge of Machine Learning, Mathematical Modeling and Operations Research. Comfortable with R, Python, SAS and Weka, MATLAB, Relational databases. Deep understanding & exposure of Data Analytics.
- Experience in Data Analytics, Machine Learning and Big Data Technologies (Hadoop, Spark, HDFS, Hive, Sqoop, PIG). Also, experience presenting data insights with Tableau dashboards and reports.
- Experience in implementing data analysis with various analytic tools, such as Anaconda 4.0 Jupiter Notebook 4.X, R 3.0 (ggplot2,, dplyr, Caret) and Excel
- Skilled in Advanced Regression Modeling, Correlation, Multivariate Analysis, Model Building, Business Intelligence tools and application of Statistical Concepts.
- Developed predictive models using Decision Tree, Naive Bayes, Logistic Regression, Random Forest, Social Network Analysis, Cluster Analysis, and Neural Networks.
- Experienced in Machine Learning and Statistical Analysis with Python Scikit - Learn.
- Extract a different kinds of raw data and perform a detailed analysis to find out a meaningful IT Services insights from it using different analytical, statistical, Machine learning techniques. E.g. ticketing sales pattern, its pricing analysis, laundry data statistics, heart patient's life style Data Science pattern, bank defaulters etc. Machine/Deep Learning Classification
- Data analysis Module design, development & implementation for the global healthcare Time Series medicine prescription data to find out of the trend, medicine usage, human health causes Predictive analysis and many other things for the further research points in a healthcare field. NLP Statistics
- Plan, develop and deploy a machine learning models to support the business. Performed Text Mining the data analysis by developing a techniques/models to sort and record a different products Pythonsizes in the system, quantity on daily/weekly/monthly basis which helped a business to R
- Hands on experience integrating R using rhdfs, hiver, rhbase packages.
- Hands on experience in SAP Hana, HDFS and R integration.
- Hands on experience in making Web application in R using shiny package. dept in Data Quality Management to get, clean, process, and cross-verify the data in multiple sources.
- Learning Machine learning course from Standford University through Coursera.
- Highly motivated team player with excellent interpersonal skills, effective communication, analytical and presentation skills.
TECHNICAL SKILLS:
Machine Learning: Regression, Polynomial Regression, Random Forest, Logistic Regression, Decision Trees, Classification, Clustering, Association, Simple/Multiple linear, Kernel SVM, K-Nearest Neighbours (K-NN).
OLAP/ BI / ETL Tool: Business Objects 6.1/XI, MS SQL Server 2008/2005 Analysis Services (MS OLAP, SSAS), Integration Services (SSIS), Reporting Services (SSRS), Performance Point Server (PPS), Oracle 9i OLAP, MS Office Web Components (OWC11), DTS, MDX, Crystal Reports 10, Crystal Enterprise 10(CMC)
Packages: ggplot2, caret, dplyr, Rweka, gmodels, RCurl, tm, C50, twitter, NLP, Reshape2, rjson, plyr, pandas, numPy, seaborn, sciPy, matplot lib, scikit-learn, Beautiful Soup, Rpy2, sqlalchemy.
Web Technologies: JDBC, HTML5, DHTML and XML, CSS3, Web Services, WSDL
Tools: Erwin r 9.6, 9.5, 9.1, 8.x, Rational Rose, ER/Studio, MS Visio, SAP Power designer.
Databases: SQL, Hive, Impala, Pig, Spark SQL, Databases SQL-Server, MySQL, MS Access, HDFS, HBase, Teradata, Netezza, Mongo DB, Cassandra, SAP HANA.
Reporting Tools: MS Office (Word/Excel/Power Point/ Visio), Tableau, Crystal reports XI, Business Intelligence, SSRS, Business Objects 5.x/ 6.x, Cognos7.0/6.0.
Operating System: Windows, Linux, Unix, Macintosh HD, Red Hat.
PROFESSIONAL EXPERIENCE:
Confidential, Seattle, WA
Data Scientist/Machine Learning Engineer
Responsibilities:
- Provided the architectural leadership in shaping strategic, business technology projects, with an emphasis on application architecture.
- Utilized domain knowledge and application portfolio knowledge to play a key role in defining the future state of large, business technology programs.
- Built models using Statistical techniques like Bayesian HMM and Machine Learning classification models like XGBoost, SVM, and Random Forest.
- Participated in all phases of data mining, data cleaning, data collection, developing models, validation, and visualization and performed Gap analysis.
- Installed and used Caffe Deep Learning Framework
- Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python.
- Worked as Data Architects and IT Architects to understand the movement of data and its storage and ER Studio 9.7
- Used pandas, numpy, seaborn, matplotlib, scikit-learn, scipy, NLTK in Python for developing various machine learning algorithms.
- Developed MapReduce/Spark Python modules for machine learning & predictive analytics on AWS. Implemented a Python-based distributed random forest via Python streaming.
- Used Pandas, NumPy, seaborn, SciPy, Matplotlib, Scikit-learn, NLTK in Python for developing various machine learning algorithms and utilized machine learning algorithms such as linear regression, multivariate regression, naive Bayes, Random Forests, K-means, & KNN for data analysis.
- Conducted studies, rapid plots and using advance data mining and statistical modelling techniques to build solution that optimize the quality and performance of data.
- Development & analysis of online ticketing aggregator application.
- Show Plazza is an online ticket aggregating system which offers a service to purchase a movies, events and programs tickets online with ticket delivery and ticket resale options. We used to get the data from different customers for the analysis of their sales. Our team used to study the behavior, trend of the customers for booking the tickets, their ps from the data and suggest the same to the client. On base of that results client would decide the types of movies, events, programs ticket they should focus for the business on priority. On the basis of analysis client would approach to their customers with the preferred events, programs they should organize for the more business according to the festive seasons with considering other parameters & factors. We were responsible for the application. development including their application enhancement, changes & support according to business needs.
- Demonstrated experience in design and implementation of Statistical models, Predictive models, enterprise data model, metadata solution and data life cycle management in both RDBMS environments.
- Analyzed large data sets apply machine learning techniques and develop predictive models, statistical models and developing and enhancing statistical models by leveraging best-in-class modeling techniques.
- Worked on database design, relational integrity constraints, OLAP, OLTP, Cubes and Normalization (3NF) and De-normalization of database.
- Worked on customer segmentation using an unsupervised learning technique - clustering.
Environment: : Erwin r9.6, Python, SQL, Oracle 12c, Netezza, SQL Server, Informatica, Java, SSRS, PL/SQL, T-SQL, Tableau, MLlib, Regression, R, Cluster analysis, Scala NLP, Spark, Kafka, MongoDB, logistic regression, Teradata, random forest, OLAP, Azure, MariaDB, SAP CRM, HDFS, ODS, NLTK, SVM, JSON, Tableau, XML, Cassandra, AWS.
Confidential, Tempe, AZ
Data Scientist
Responsibilities:
- Extracted data from HDFS and prepared data for exploratory analysis using data munging
- Built models using Statistical techniques like Bayesian HMM and Machine Learning classification models like XGBoost, SVM, and Random Forest.
- Involved in discussion with business people to understand the requirements and decided which category of data needed for and testing the ML model.
- Involved in collecting data from different cross functional businesses (live data) and stored in appropriate data stores (Relation/NoSQLdatabases, different file formats) by using data pipelining methods.
- Transformed raw data into statistical features and applied most efficient Machine learning models (predictive, recommendation, regression, classification, and forecasting and ensemble models) to achieve the high accuracy model.
- Determined important KPIs needed for the MLmodels and achieved it with high accuracy by conducting various metrics techinques (Confusion matrix, Classification accuracy, ROC, precision, etc., ).
- Did A/B and hypothesis testing for MLmodels to ensure the accuracy and efficiency of the model.
- Developed dashboards using Tableau.
- Built various in-house products (predictive, recommendation and NLP models) which reduced the man power and cost for the client.
- Built an image processing deep learning model which reduce the insurance claim time for the customer and eased the document process.
- Experience in developing programs in Spark using Python to compare the performance of Spark with Hive and SQL/Oracle.
- Data Manipulation and Aggregation from different source using Nexus, Business Objects, Toad, Power BI and Smart View.
- Experience in handling multiple relational databases like SQL Server, Oracle
- Focus on integration overlap and Informatica newer commitment to MDM with the acquisition of Identity Systems.
- Good knowledge on Spark components like Spark SQL, MLib, Spark Streaming and GraphX,
- Extensively worked on Spark Streaming and Apache Kafka to fetch live stream data.
- Implemented novel algorithm for test and control team using Spark /Scala, Oozie, HDFS and Python on P&G Yarn cluster.
- Worked on Python Open stack API's.
- Skilled in using dplyr and pandas in R and Python for performing exploratory data analysis.
- Developed scalable model using Spark (RDD, Mllib, Ml, Dataframes) in Scala
- Integrated Tesseract, ghost script with Spark to access data in hdfs and saving data in hive table
- Coded proprietary packages to analyze and visualize SPCfile data to identify bad spectra and samples to reduce unnecessary procedures and costs.
- Implemented Classification using supervised algorithms like Logistic Regression, Decision trees, Naive Bayes, KNN.
- Updated Python scripts to match data with our database stored in AWS Cloud Search, so that we would be able to assign each document a response label for further classification.
- Data transformation from various resources, data organization, features extraction from raw and stored.
- Validated the machine learning classifiers using ROC Curves and Lift Charts.
Environment: Unix, Python 3.5.2, MLLib, SAS, regression, logistic regression,, NoSQL, Teradata, OLTP, random forest, OLAP, HDFS, ODS, NLTK, SVM, JSON, XML and MapReduce.
Confidential, Southlake, TX
Data Scientist
Responsibilities:
- Identified areas of improvement in existing business by unearthing insights by analysing vast amount of data using machine learning techniques.
- Interpret problems and provides solutions to business problems using data analysis, data mining, optimization tools, and machine learning techniques and statistics.
- Led technical implementation of advanced analytics projects, Defined the mathematical approaches, developer new and effective analytics algorithms and wrote the key pieces of mission-critical source code implementing advanced machine learning algorithms utilizing caffe, TensorFlow, Scala, Spark, MLLib, R and other tools and languages needed.
- Gathers, analyses, documents and translates application requirements into data models and Supports standardization of documentation and the adoption of standards and practices related to data and applications.
- Designed and developed NLP models for sentiment analysis.
- Led discussions with users to gather business processes requirements and data requirements to develop a variety of Conceptual, Logical and Physical Data Models. Expert in Business Intelligence and Data Visualization tools: Tableau, Microstrategy.
- Designed and implemented system architecture for Confidential EC2 based cloud-hosted solution for client.
- Designed the Enterprise Conceptual, Logical, and Physical Data Model for 'Bulk Data Storage System 'using Embarcadero ER Studio, the data models were designed in 3NF
- Worked on machine learning on large size data using Spark and MapReduce.
Environment: Python,PySpark, Tableau, MongoDB, SQL Server, SDLC, ETL, SSIS, recommendation systems, Machine Learning Algorithms, text-mining process, A/B test.
Confidential, Dallas, TX
Data Scientist
Responsibilities:
- Worked with BI team in gathering the report requirements & also Sqoop to export data into HDFS & Hive
- Involved in the below phases of Analytics using R, Python and Jupyter notebook.
- Data collection and treatment: Analysed existing internal data and external data, worked on entry errors, classification errors and defined criteria for missing values.
- Used cluster analysis for identifying customer segments, Decision trees used for profitable and non-profitable customers, Market Basket Analysis used for customer purchasing behaviour and part/product association.
- Developed multiple Map Reduce jobs in Java for data cleaning and preprocessing.
- Assisted with data capacity planning and node forecasting.
- Installed, Configured and managed Flume Infrastructure
- Administrator for Pig, Hive and HBase installing updates patches and upgrades.
- Worked closely with the claims processing team to obtain patterns in filing of fraudulent claims.
- Worked on performing major upgrade of cluster from CDH3u6 to CDH4.4.0
- Developed Map Reduce programs to extract and transform the data sets and results were exported back to RDBMS using Sqoop.
- Patterns were observed in fraudulent claims using text mining in R and Hive.
- Exported the data required information to RDBMS using Sqoop to make the data available for the claims processing team to assist in processing a claim based on the data.
- Developed Map Reduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.
Environment: HDFS, PIG, HIVE, Map Reduce, Linux, HBase, Flume, Sqoop, R, VMware, Eclipse, C
Confidential
Data Analyst
Responsibilities:
- Worked with BI team in gathering the report requirements and also Sqoop to export data into HDFS and Hive
- Involved in the below phases of Analytics using R, Python and Jupyter notebook.
- Data collection and treatment: Analysed existing internal data and external data, worked on entry errors,classification errors and defined criteria for missing values
- Assisted with data capacity planning and node forecasting.
- Installed, Configured and managed Flume Infrastructure .
- Worked closely with the claims processing team to obtain patterns in filing of fraudulent claims.
- Worked on performing major upgrade of cluster from CDH3u6 to CDH4.4.0
- Developed Map Reduce programs to extract and transform the data sets and results were exported back to RDBMS using Sqoop.
- Patterns were observed in fraudulent claims using text mining in R and Hive.
- Exported the data required information to RDBMS using Sqoop to make the data available for the claims processing team to assist in processing a claim based on the data.
- Developed Map Reduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.
- Created tables in Hive and loaded the structured (resulted from Map Reduce jobs) data
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with EDW tables and historical metrics.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
- Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems.
- Managed and reviewed Hadoop log files.
- Tested raw data and executed performance scripts.
Environment: HDFS, PIG, HIVE, Map Reduce, Linux, HBase, Flume, Sqoop, R, VMware, Eclipse, Cloudera, Python.
Confidential
Data Analyst
Responsibilities:
- Wrote SQL queries for data validation on the backend systems and used various tools like TOAD&DBVisualizer for DBMS(Oracle)
- Perform Data analysis, Backend Database testing, Data Modeling and Developing SQL Queries to solve problems and meet user's need for Database management in Data Warehouse.
- Utilize object-oriented languages, concepts, database design, star schemas and databases.
- Create algorithms as needed to manage and implement proposed solutions.
- Participate in test planning and test execution for functional, system, integration, regression, UAT (User Acceptance Testing), load and performance testing.
- Work with test automation tools for recording/coding in Database, and execute in regression testing cycles.
- Transferred data from various OLTP data sources, such as Oracle, MS Access, MS Excel, Flat files, CSV files into SQL Server.
- Working with Databases DB2, Oracle DM, SQL Server for Database testing and maintenance.
- Involved in writing and executing User Acceptance Testing (UAT) with end users.
- Involved in Post- Implementation validations after the changes have been to the Data Marts.
- Chart out Graphs, and Reports alike in QC to point out the percentage of Test Cases passed, and thereby to point out the percentage of Quality achieved and uploading the status daily to ART reports an in-house tool.
- Performed extensive Data Validation, Data Verification against Data Warehouse.
- Used UNIX to check the Data marts, Tables and Updates made to the tables.
- Writing advanced SQL Queries to query the data from Data marts and Landings to verify the changes has been made.
- Involved in Client requirement gathering, participated in discussion & brain storming sessions and documented requirements.
- V alidating and profilingFlat File Data into Teradata tables using UNIX Shell scripts.
- Actively participated Functional, System and User Acceptance testing on all builds and supervised releases to ensure system / functionality integrity.
- Interacted with developers to resolve different Quality Related Issues.
- Wrote and executed manual test cases for functional, GUI, and regression testing of the application to make sure that new enhancements do not break working features
- Writing and executing Manual test cases in HP Quality Center.
- Wrote test plans for positive and negative scenarios for GUI and functional testing
- Involved in writing SQL queries and stored procedures using Query Analyzer and matched the results retrieved from the batch log files
- Created Project Charter documents & Detailed Requirement document and reviewed with Development & other stake holders.
Environment: Subversion, Tortoise SVN, Jira, Agile-Scrum, Web Services, Mainframe, Oracle, Perl, UNIX, LINUX, Shell Scripts, UML, Quality Center,, RequisitePro, SQL, MS Visio, MS Project, Excel, Power Point, Word, SharePoint, Win XP/7 Enterprise.
