Data Scientist Resume
Mountainview, CA
SUMMARY
- Above 8+ years of experience in Machine Learning, Data Mining with large datasets of Structured and Unstructured data, Data Acquisition, DataValidation, Predictive Modeling, Data Visualization.
- Extensive experience in Text Analytics, developing different Statistical Machine Learning, Data Mining solutions to various business problems and generating data visualizations using R, Python and Tableau.
- Designing of Physical Data Architecture of New system engines.
- Hands on SparkMlib utilities such as including classification, regression, clustering, collaborative filtering, dimensionality reduction.
- Proficient in Predictive Modeling,DataMining Methods, Factor Analysis, ANOVA, Hypothetical testing, normal distribution and other advanced statistical and econometric techniques.
- Developed predictive models using Decision Tree, RandomForest, Naïve Bayes, Logistic Regression, Cluster Analysis, and Neural Networks.
- Experienced the full software life cycle in SDLC, Agile and Scrum methodologies.
- Skilled in Advanced Regression Modeling, Correlation, Multivariate Analysis, Model Building, Business Intelligence tools and application of Statistical Concepts.
- Experienced in Machine Learning and Statistical Analysis with Python Scikit - Learn.
- Strong SQL programming skills, with experience in working with functions, packages and triggers.
- Experienced in Visual Basic for Applications and VB programming languages to work with developing applications.
- Excellent understanding of machine learning techniques and algorithms, such as k-NN, Naive Bayes, SVM, Decision Forests, natural language processing (NLP) etc.
- Experienced in Python to manipulatedatafordataloading and extraction and worked with python libraries like Matplotlib, Numpy, Scipy and Pandas fordataanalysis.
- Worked with complex applications such as R, Scala, SAS, Matlab and SPSS to develop neural network, cluster analysis.
- Worked with RDBMS including MySQL, DB2 and Oracle SQL.
- Worked with NoSQL Database including Hbase, Cassandra and MongoDB.
- Experienced inDataIntegration Validation andDataQuality controls for ETL process andData Warehousing using MS Visual Studio SSIS, SSAS, SSRS.
- Understanding ofSAS andR programming, conducting data extracting analysis,profiling and composite reporttools- SAS/R/SQL/ PYTHON / TABLEAUto perform joinsExpertise in data facts collection, root cause analysis of business problem and formulation of the key findings to the management.
- Highly skilled in using Hadoop (pig and Hive) for basic analysis and extraction of data in the infrastructure to provide Data Summarization.
- Highly skilled in using Visualization tools like Tableau, ggplot2 and d3.js for creating dashboards.
- Worked and extracteddatafrom various database sources like Oracle, SQL Server, DB2, Regularly accessing JIRA tool and other internal issue trackers for the Project development.
- Skilled in System Analysis, E-R/DimensionalDataModeling, Database Design and implementing RDBMS specific features.
- Knowledge of working with Proof of Concepts (PoC’s) and gap analysis and gathered necessary data for analysis from different sources, prepared data for data exploration using data munging and Teradata.
- Solid coding and engineering skills preferably in Machine Learning.
TECHNICAL SKILLS
Languages: T-SQL, PL/SQL, SQL, C, C++, XML, HTML, DHTML, HTTP, Matlab, DAX, Python
Databases: SQL Server 2014/2012/2008/2005/2000 , MS-Access, Oracle 11g/10g/9i and Teradata, big data, Hadoop
BI Tools: Microsoft Power BI, Tableau, SSIS, SSRS, SSAS, Business Intelligence Development Studio (BIDS), Visual Studio, Crystal Reports, Informatica 6.1.
Database Design Tools and Data Modeling: MS Visio, ERWIN 4.5/4.0, Star Schema/Snowflake Schema modeling, Fact & Dimensions tables, physical & logical data modeling, Normalization and De-normalization techniques, Kimball &Inmon Methodologies
Tools: and Utilities: SQL Server Management Studio, SQL Server Enterprise Manager, SQL Server Profiler, Import & Export Wizard, Visual Studio .Net, Microsoft Management Console, Visual Source Safe 6.0, DTS, Crystal Reports, Power Pivot, ProClarity, Microsoft Office, Excel Power Pivot, Excel Data Explorer, Tableau, JIRA, SparkMlib
PROFESSIONAL EXPERIENCE
DATA SCIENTIST
Confidential, MOUNTAINVIEW, CA
Responsibilities:
- Created Hive and Tez views and providing permission to user and AD groups on Ambari server.
- Created Hive internal/external tables with proper static and dynamic partitions and working on them using HQL.
- Deployed data from various sources into HDFS and building reports using Tableau.
- Written Hive queries for data analysis to meet the business requirement.
- Performance tuning using Partitioning, bucketing of HIVE tables.
- Worked on creating UDF for Hive and Impala.
- Designing, building, installing, configuring and supporting Hadoop.
- Developed MapReduce programs in Java for data cleansing and pre-processing.
- Worked on creating the RDD's, Data Frame's for the required input data and performed the data transformations using Spark Scala.
- Import the data from different sources like HDFS/Hbase into Spark RDD.
- Developed spark scripts by using Scala shell as per requirements.
- Used spark cluster to manipulate RDDS (Resilient Distributed Datasets) and also used concepts of RDD partitions.
- Handled importing data from different data sources into HDFS using Sqoop and also performing transformations using Hive, MapReduce and then loading data into HDFS.
- Used the Spark - Cassandra Connector to load data to and from Cassandra. Real time streaming the data using Spark with .
- Loading data into spark RDD and do in memory data Computation to generate the Output response.
- Integrated Kafka and Spark Streaming using Scala.
- Created SQL tables with referential integrity and developed queries using SQL, SQL*PLUS and PL/SQL.
- Formulated procedures for integration of R programming plans with data sources and delivery systems and R language was used for prediction.
- Implementing SparkMlib utilities such as including classification, regression, clustering, collaborative filtering and dimensionality reduction.
- Design, coding, unit testing of ETL package source marts and subject marts using Informatica ETL processes for Oracledatabase.
- Developed Statistical Analysis and Response Modeling for Analytical Data base contributors (logistic regression).
- Developed various QlikView Data Models by extracting and using the data from various sources files, DB2, Excel, Flat Files and Bigdata.
- Handled importing data from various data sources, performed transformations using Hive, Map Reduce, and loaded data into HDFS.
- Interaction with Business Analyst, SMEs and other Data Architects to understand Business needs and functionality for various project solutions
- Researched, evaluated, architected, and deployed new tools, frameworks, and patterns to built sustainable platforms for the clients
- Identifying and executing process improvements, hands-on in various technologies such as Oracle, Informatica, Business Objects.
- Designed both 3NF data models for ODS, OLTP systems and dimensional data models using Star and Snowflake Schemas.
Environment: r9.0, Informatica 9.0, ODS, OLTP, Oracle 10g, Hive, OLAP, DB2, Metadata, MS Excel, Mainframes MS Visio, Rational Rose, Requisite Pro., Hadoop, PL/SQL, SAS etc.
DATA SCIENTIST
Confidential, BIRMINGHAM, AL.
Responsibilities:
- Worked as aDataModeler/Analyst to generateDataModels using Erwin and developed relational database system.
- Analyzed the business requirements of the project by studying the Business Requirement Specification document.
- Participated in installation ofSAS/EBI on LINUX platform
- Extensively worked on Data Modeling tools Erwin Data Modeler to design the data models.
- Designed mapping to process the incremental changes that exists in the source table. Whenever source data elements were missing in source tables, these were modified/added in consistency with third normal form based OLTP source database.
- Designed tables and implemented the naming conventions for Logical and Physical Data Models in Erwin 7.0.
- Explored and Extracteddatafrom source XML in HDFS, preparingdatafor exploratory analysis usingdatamunging.
- Used R and Python for ExploratoryDataAnalysis, A/B testing, Anova test and Hypothesis test to compare and identify the effectiveness of Creative Campaigns.
- Used Spark for testdataanalytics using MLLib and Analyzed the performance to identify bottlenecks.
- Worked on differentdataformats such as JSON, XML and performed machine learning algorithms in R.
- Performed performance improvement of the existingDatawarehouse applications to increase efficiency of the existing system.
- Designed and developed Use Case, Activity Diagrams, Sequence Diagrams, OOD (Object oriented Design) using UML and Vision.
- Worked on Linux shell scripts for business process and loadingdatafrom different interfaces to HDFS.
- Created MDM, OLAPdataarchitecture, analyticaldatamarts, and cubes optimized for reporting.
- Worked with different sources such as Oracle, Teradata, SQL Server2012 and Excel, Flat, Complex Flat File, Cassandra, MongoDB, HBase, and COBOL files.
- Performed K-means clustering, Multivariate analysis and Support Vector Machines in R.
- Used Python, R, SQL to create Statistical algorithms involving Multivariate Regression, Linear Regression, Logistic Regression, PCA, Random forest models, Decision trees, Support Vector Machine for estimating the risks of welfare dependency.
- Identified and targeted welfare high-risk groups with Machine learning algorithms.
- Created PL/SQL packages and Database Triggers and developed user procedures and prepared user manuals for the new programs.
Environment: SQL Server 2008R2 / 2005 Enterprise, SSRS, SSIS, Crystal Reports, Windows Enterprise Server 2000, DTS, SQL Profiler, and Query Analyzer. Python, MDM, MLLib, PL/SQL, Tableau, Teradata 14.1, JSON, HADOOP (HDFS), MapReduce, SQL Server, MLLib, Scala NLP, SSMS, ERP, CRM, Netezza, Cassandra, SQL, PL/SQL, SSRS, Informatica, PIG, Spark, Azure, R Studio, MongoDB, JAVA, HIVE.
DATA SCIENTIST
Confidential, DALLAS, TX
Responsibilities:
- Coded R functions to interface with Caffe Deep Learning Framework.
- Working in Amazon Web Services cloud computing environment
- Worked with several R packages including knitr, dplyr, SparkR, CausalInfer, spacetime.
- Implemented end-to-end systems for Data Analytics, Data Automation and integrated with custom visualization tools using R, Mahout, Hadoop and MongoDB.
- Gathering all the data that is required from multiple data sources and creating datasets that will be used in analysis.
- Performed Exploratory Data Analysis and Data Visualizations using R and Tableau.
- Perform a proper EDA, Univariate and bi-variate analysis to understand the intrinsic effect/combined effects.
- Worked withDatagovernance,Dataquality,datalineage,Dataarchitectto design various models and processes.
- Independently coded new programs and designed Tables to load and test the program effectively for the given POC's using with BigData/Hadoop.
- Designeddatamodels anddataflow diagrams using Erwin and MSVisio.
- As an Architect implemented MDM hub to provide clean, consistent data for a SOA implementation.
- Developed, Implemented & Maintained the Conceptual, Logical&PhysicalDataModels using Erwin for Forward/Reverse Engineered Databases.
- EstablishedDataarchitecture strategy, best practices, standards, and roadmaps.
- Lead the development and presentation of a data analytics data-hub prototype with the help of the other members of the emerging solutions team
- Performed data cleaning and imputation of missing values using R.
- Worked with Hadoop eco system covering HDFS, HBase, YARN and MapReduce.
- Take up ad-hoc requests based on different departments and locations
- Used Hive to store the data and perform data cleaning steps for huge datasets.
- Created dash boards and visualization on regular basis using ggplot2 and Tableau.
- Creating customized business reports and sharing insights to the management.
- Worked with BTEQ to submit SQL statements, import and export data, and generate reports in Teradata.
- Interacted with the other departments to understand and identify data needs and requirements and work with other members of the IT organization to deliver data visualization and reporting solutions to address those needs.
Environment: Erwin r9.0, Informatica 9.0, ODS, OLTP, Oracle 10g, Hive, OLAP, DB2, Metadata, MS Excel, Mainframes MS Visio, Rational Rose, Requisite Pro, Hadoop, PL/SQL, etc.
DATA ANALYST
Confidential, MINNEAPOLIS, MN
Responsibilities:
- Gathering the requirements by interacting heavily with the business users, multiple technical teams to design and develop the workflows for the new functional piece.
- Collaborated with various business stakeholders to create Business Requirement Document (BRD), translated gathered high-level requirements into a Functional Requirement Document (FRD) to assist implementation side SMEs and developers, along withdataflow diagrams, user stories and use casesPart of a Scrum Agile team.
- Experience in SQL joins, sub queries, tracing and performance tuning for better running of queries
- Extensively used joins and sub queries for complex queries involving multiple tables from different databases.
- Performance tuning of stored procedures and functions to optimize the query for better performance.
- Successfully implemented indexes on tables for optimum performance.
- Developed complex stored procedures using T-SQL to generate Ad-hoc reports within SQL Server. Reporting services.
- Strong analytical, problem solving skills coupled with interpersonal, and leadership skills.
- Interaction with the clients to gather out the requirements and assist them in immediate workarounds for the issues with the application.
- Developed detailed test scenarios as documented in business requirements documents, assisted the test team with UAT.
- Real time usage of Tableau for analytical purpose.
Environment: SQL/Server, Oracle 9i, MS-Office, Cognos, Teradata, Informatica, ER Studio, XML, Business Objects. MY SQL, MS Power Point, MS Access, T- SQL, MS Power Point, MS Access, T-SQL, DTS, SSIS, SSRS, SSAS, ETL, MDM, Teradata.
DATA ANALYST
Confidential
Responsibilities:
- Deployed GUI pages by using JSP, JSTL, HTML, DHTML, XHTML, CSS, JavaScript, AJAX.
- Configured the project on WebSphere 6.1 application servers
- Implemented the online application by using Core Java, JDBC, JSP, Servlets and EJB 1.1, Web Services, SOAP, WSDL.
- Communicated with other Health Care info by using Web Services with the help of SOAP, WSDL JAX-RPC.
- Used Singleton, factory design pattern, DAO Design Patterns based on the application requirements
- Used SAX and DOM parsers to parse the raw XML documents
- Used RAD as Development IDE for web applications.
- Create ETL scripts using Regular Expressions and custom tools (Informatica, Pentaho, and Sync Sort) to ETL data.
- Developed SQL Service Broker to flow and sync of data from MS-I to Microsoft's master database management (MDM).
- Involved in loading data between Netezza tables using NZSQL utility.
- Worked on Data Modeling using Dimensional Data Modeling, Star Schema/Snow Flake schema, and Fact & Dimensional, Physical & Logical Data Modeling.
- Generated Stats pack/AWR reports from Oracle database and analyzed the reports for Oracle8.x wait events, time consuming SQL queries, table space growth, and database growth.
- Preparing and executing Unit test cases.
- Used Log4J logging framework to write Log messages with various levels.
- Involved in fixing bugs and minor enhancements for the front-end modules.
- Implemented Microsoft Visio and Rational Rose for designing the Use Case Diagrams, Class model, Sequence diagrams, and Activity diagrams for SDLC process of the application.
- Doing functional and technical reviews.
- Maintenance in the testing team for System testing/Integration/UAT.
- Guaranteeing quality in the deliverables.
- Conducted Design reviews and Technical reviews with other project stakeholders.
- Was a part of the complete life cycle of the project from the requirements to the production support.
- Created test plan documents for all back-end database modules.
- Implemented the project in Linux environment.
Environment: R 3.0, Erwin 9.5, Tableau 8.0, MDM, QlikView, MLLib, PL/SQL, HDFS, Teradata 14.1, JSON, HADOOP (HDFS), MapReduce, PIG, Spark, R Studio, MAHOUT, JAVA, HIVE, AWS.
DATA ANALYST / DATA MODELER
Confidential
Responsibilities:
- Worked with project team representatives to ensure that logical and physical ER/Studio data models were developed in line with corporate standards and guidelines.
- Involved in defining the source to target data mappings, business rules, data definitions.
- Worked with BTEQ to submitSQLstatements, import and export data, and generate reports in Teradata.
- Responsible for defining the key identifiers for each mapping/interface.
- Responsible for defining the functional requirement documents for each source to target interface.
- Document, clarify, and communicate requests for change requests with the requestor and coordinate with the development and testing team.
- Work with users to identify the most appropriate source of record and profile the data required for sales and service.
- Implementation of Metadata Repository, Maintaining Data Quality, Data Clean up procedures, Transformations, Data Standards, Data Governance program, Scripts, Stored Procedures, triggers and execution of test plans
- Performed data quality in Talend Open Studio.
- Coordinated meetings with vendors to define requirements and system interaction agreement documentation between client and vendor system.
- Document data quality and traceability documents for each source interface.
- Generate weekly and monthly asset inventory reports.
Environment: ER Studio, MY SQL, MS Power Point, MS Access, MY SQL, MS Power Point, MS Access, Netezza, DB2, T-SQL, DTS, SSIS, SSRS, SSAS, ETL, MDM, Teradata, Oracle8.x, (Star Schema and Snow Flake Schema) etc.
