We provide IT Staff Augmentation Services!

Data Scientist Resume

0/5 (Submit Your Rating)

Los Angeles, CA

SUMMARY

  • Above 8+ years of experience in MachineLearning, Datamining with largedatasets of Structured and Unstructureddata, DataAcquisition, DataValidation, Predictivemodeling, DataVisualization.
  • Extensive experience in Text Analytics, developing different StatisticalMachineLearning, DataMining solutions to various business problems and generating datavisualizations using R, Python and Tableau.
  • Designing of PhysicalDataArchitecture of New system engines.
  • Hands on experience in implementing LDA, NaiveBayes and skilled in RandomForests, Decision Trees, Linear and Logistic Regression, SVM, Clustering, neuralnetworks, Principal Component Analysis and good knowledge on Recommender Systems.
  • Proficient in StatisticalModeling and MachineLearning techniques (Linear, Logistics, Decision Trees, Random Forest, SVM, K - Nearest Neighbors, Bayesian, XG Boost) in Forecasting/ Predictive Analytics, Segmentation methodologies, Regression based models, Hypothesis testing, Factor analysis/ PCA, Ensembles.
  • Expertise in transforming business requirements into analyticalmodels, designingalgorithms, buildingmodels, developing datamining and reportingsolutions dat scales across massive volume of structured and unstructured data.
  • Developing LogicalDataArchitecture with adherence to Enterprise Architecture.
  • Strong experience in Software Development Life Cycle (SDLC) including RequirementsAnalysis, Design Specification and Testing as per Cycle in both Waterfall and Agile methodologies.
  • Adept in statisticalprogramminglanguages like Rand alsoPython including BigData technologies like Hadoop, Hive.
  • Skilled in using dplyr and pandas in R and python for performing Exploratory dataanalysis.
  • Experience working withdatamodeling tools like Erwin, PowerDesigner and ERStudio.
  • Experience in designing starschema, Snowflakeschema forDataWarehouse, ODSarchitecture.
  • Experience in designing stunning visualizations using Tableau software and publishing and presenting dashboards, Storyline on web and desktop platforms.
  • Experience and Technicalproficiency in Designing,DataModelingOnlineApplications, Solution Lead for ArchitectingDataWarehouse/BusinessIntelligence Applications.
  • Good understanding of TeradataSQLAssistant, Teradata Administrator anddataload/ export utilities like BTEQ, FastLoad, MultiLoad, FastExport.
  • Experience with DataAnalytics, DataReporting, Ad-hocReporting,Graphs, Scales, PivotTables and OLAP reporting.
  • Highly skilled in using Hadoop (pig and Hive) for basic analysis and extraction of data in the infrastructure to provide datasummarization.
  • Highly skilled in using visualization tools like Tableau, ggplot2 and d3.js for creating dashboards.
  • Worked and extracteddatafrom various database sources like Oracle, SQLServer, DB2, Regularly accessing JIRA tool and other internal issue trackers for the Project development.
  • Skilled in System Analysis, E-R/DimensionalDataModeling, DatabaseDesign and implementing RDBMS specific features.
  • Knowledge of working with Proof of Concepts (PoC’s) and gapanalysis and gatheird necessary data for analysis from different sources, prepared data for dataexploration using datamunging and Teradata.
  • Well experienced in Normalization&De-Normalization techniques for optimum performance in relational and dimensional database environments

TECHNICAL SKILLS

Data Modeling Tools: Erwin r9.6/9.5, ER/Studio 9.7, Star-Schema Modeling, Snowflake-Schema Modeling, FACT and dimension tables, Pivot Tables.

Frameworks: Microsoft .Net 4.5/ 4.0/ 3.5/3.0 , Entity Framework, Bootstrap, Microsoft Azure, Swagger.

Database Design Tools and Data Modeling: MS Visio, ERWIN 4.5/4.0, Star Schema/Snowflake Schema modeling, Fact & Dimensions tables, physical & logical data modeling, Normalization and De-normalization techniques, Kimball &Inmon Methodologies

Version Controller: TFS, Microsoft Visual SourceSafe, GIT, NUNIT, MSUNIT

Software Packages: MS-Office 2003/ 07/10/13 , MS Access, Messaging Architectures.

Operating Systems: Windows Win8/XP/NT/ 95/98/2000/2008/2012 , Android SDK.

Microsoft Technologies: PHP,Scala2,Shark2,Awk,Cascading,Cassandra,Clojure,Fortran,JavaScript,JMP,Mahout,objectiveC,QlickView,Redis,Redshifed

Web Technologies: Windows API, Web Services, Web API (RESTFUL) HTML5, XHTML, CSS3, AJAX, XML, XAML,MSMQ, Silverlight, Kendo UI.

DWH / BI Tools: Microsoft Power BI, Tableau, SSIS, SSRS, SSAS, Business Intelligence Development Studio (BIDS), Visual Studio, Crystal Reports, Informatica 6.1.

Programming Languages: C#, VB.NET (VB6), VBScript, OOPS, Data structures, Matlab,Algorithms, Python, R, Java, Java Script, SQL, J2EE, C, C++ and XML.

Development Tools: R x 30, SQL x 27, Python x 22, Hadoop x 19, SAS x 18, Java x15, Hive x 13, Mat lab x 12, R Studio,SAS, MSOffice,Visual Studio 2010

Big Data Tools: Hadoop 2.7.2, Hive, Spark2.1.1, Pig, HBase, Sqoop, Flume

PROFESSIONAL EXPERIENCE

Confidential, Los Angeles, CA

Data Scientist

Responsibilities:

  • As an Architect design conceptual, logical and physical models using Erwin and build datamarts using hybrid Inmon and Kimball DW methodologies.
  • Worked closely with business, datagovernance, SMEs and vendors to define data requirements.
  • Worked with data investigation, discovery and mapping tools to scan every single data record from many sources.
  • Perform data integrity checks, data cleaning, exploratory analysis and feature engineer using R 3.4.0.
  • Designed the prototype of theDatamart and documented possible outcome from it for end-user.
  • Involved in business process modeling using UML
  • Visualize data in R using various packages like ggplot2 and various regression techniques
  • Developed and maintaineddatadictionary to create metadata reports for technical and business purpose.
  • Created SQLtables with referential integrity and developed queries using SQL, SQL*PLUS and PL/SQL
  • Experience in maintaining database architecture and metadata dat support the EnterpriseDatawarehouse.
  • Design, coding, unit testing of ETL package source marts and subject marts using Informatica ETL processes for Oracledatabase.
  • Developed various QlikView Data Models by extracting and using the data from various sources files, DB2, Excel, Flat Files and Bigdata.
  • Handled importing data from various data sources, performed transformations using Hive, Map Reduce, and loaded data into HDFS.
  • Interaction with Business Analyst, SMEs and other Data Architects to understand Business needs and functionality for various project solutions
  • Researched, evaluated, architected, and deployed new tools, frameworks, and patterns to built sustainable Big Data platforms for the clients
  • Identifying and executing process improvements, hands-on in various technologies such as Oracle, Informatica, BusinessObjects.
  • Designed both 3NF data models for ODS, OLTP systems and dimensionaldatamodels using Star and SnowflakeSchemas.

Environment: r9.0, Informatica 9.0, ODS, R, OLTP, Oracle 10g, Hive, OLAP, DB2, Metadata, MS Excel, Mainframes MS Visio, Rational Rose, Requisite Pro. Hadoop, PL/SQL, etc.

Confidential, Des Moines, IA

Data Scientist

Responsibilities:

  • Worked as aDataModeler/Analyst to generateDataModels using Erwin and developed relational database system.
  • Analyzed the business requirements of the project by studying the Business Requirement Specification document.
  • Extensively worked on DataModeling tools ErwinDataModeler to design the datamodels.
  • Designedmapping to process the incremental changes dat exists in the source table. Whenever source data elements were missing in source tables, these were modified/added in consistency with third normal form based OLTP source database.
  • Designed tables and implemented the naming conventions for Logical and PhysicalData Models in Erwin 7.0.
  • Provide expertise and recommendations for physicaldatabasedesign,architecture, testing, performance tuning and implementation.
  • Designedlogical and physical data models for multiple OLTP and Analytic applications.
  • Extensively used the Erwin design tool &Erwin model manager to create and maintain the DataMart.
  • Designed the physical model for implementing the model into oracle9i physicaldatabase.
  • Involved with DataAnalysis primarily Identifying DataSets, SourceData, Source Meta Data, Data Definitions and Data Formats
  • Performance tuning of the database, which includes indexes, and optimizing SQL statements, monitoring the server.
  • Wrote simple and advanced SQLqueries and scripts to create standard and adhoc reports for senior managers.
  • Collaborated thedatamapping document from source to target and thedataquality assessments for the sourcedata.
  • Used Expert level understanding of different databases in combinations for Data extraction and loading, joiningdata extracted from different databases and loading to a specific database.
  • Co-ordinate with various business users, stakeholders and SME to get Functional expertise, design and business test scenarios review, UAT participation and validation of financial data.
  • Worked very close withDataArchitectsand DBA team to implementdatamodel changes in database in all environments.
  • Created PL/SQL packages and DatabaseTriggers and developed user procedures and prepared user manuals for the new programs.
  • Performed performance improvement of the existingDatawarehouse applications to increase efficiency of the existing system.
  • Designed and developed UseCase, Activity Diagrams, Sequence Diagrams, OOD (Object oriented Design) using UML and Visio.

Environment: SQL Server 2008R2/2005 Enterprise, SSRS, SSIS, Crystal Reports, Windows Enterprise Server 2000, DTS, SQL Profiler, and Query Analyzer.

Confidential, New Jersey

Data Scientist

Responsibilities:

  • Coded R functions to interface with CaffeDeepLearning Framework
  • Working in AmazonWebServices cloud computing environment
  • Worked with several R packages including knitr, dplyr, SparkR, CausalInfer, spacetime.
  • Implemented end-to-end systems for DataAnalytics, DataAutomation and integrated with custom visualization tools using R,Mahout, Hadoop and MongoDB.
  • Gathering all the data dat is required from multiple data sources and creating datasets dat will be used in analysis.
  • Performed Exploratory DataAnalysis and DataVisualizations using R, andTableau.
  • Perform a proper EDA, Univariate and bi-variate analysis to understand the intrinsic effect/combined effects.
  • Worked withDatagovernance,Dataquality,datalineage,Dataarchitectto design various models and processes.
  • Independently coded new programs and designed Tables to load and test the program effectively for the given POC's using with BigData/Hadoop.
  • Designeddatamodels anddataflow diagrams using Erwin and MSVisio.
  • As an Architect implemented MDM hub to provide clean, consistent data for a SOA implementation.
  • Developed, Implemented & Maintained the Conceptual, Logical&PhysicalDataModels using Erwin for Forward/ReverseEngineered Databases.
  • EstablishedDataarchitecture strategy, best practices, standards, and roadmaps.
  • Lead the development and presentation of a dataanalytics data-hub prototype with the halp of the other members of the emerging solutions team
  • Performed datacleaning and imputation of missing values using R.
  • Worked with Hadoop eco system covering HDFS, HBase, YARN and MapReduce
  • Take up ad-hoc requests based on different departments and locations
  • Used Hive to store the data and perform datacleaning steps for huge datasets.
  • Created dash boards and visualization on regular basis using ggplot2 and Tableau
  • Creating customized business reports and sharing insights to the management
  • Worked with BTEQ to submit SQL statements, import and export data, and generate reports in Teradata.
  • Interacted with the other departments to understand and identify dataneeds and requirements and work with other members of the ITorganization to deliver data visualization and reportingsolutions to address those needs.

Environment: Erwin r9.0, Informatica 9.0, ODS, OLTP, Oracle 10g, Hive, OLAP, DB2, Metadata, MS Excel, Mainframes MS Visio, Rational Rose, Requisite Pro. Hadoop, PL/SQL, etc..

Confidential, Indianapolis, IN

Data Analyst

Responsibilities:

  • Supported MapReduce Programs running on the cluster.
  • Evaluated business requirements and prepared detailed specifications dat follow project guidelines required to develop written programs.
  • Configured Hadoop cluster with Namenode and slaves and formatted HDFS.
  • Used Oozieworkflow engine to run multiple Hive and Pig jobs.
  • Performed Map Reduce Programs those are running on the cluster.
  • Developed multiple MapReduce jobs in java for data cleaning and preprocessing.
  • Analyzed the partitioned and bucketed data and compute various metrics for reporting.
  • Involved in loading data from RDBMS and web logs into HDFS using SqoopandFlume.
  • Worked on loading the data from MySQL to HBase where necessary using Sqoop.
  • Developed Hive queries for Analysis across different banners.
  • Extracted data from Twitter using Java and Twitter API. Parsed JSON formatted twitter data and uploaded to database.
  • Launching Amazon EC2 Cloud Instances using Amazon Images (Linux/ Ubuntu) and Configuring launched instances with respect to specific applications.
  • Exported the result set from Hive to MySQL using Sqoop after processing the data.
  • Analyzed the data by performing Hive queries and running Pig scripts to study customer behavior.
  • Have hands on experience working on Sequence files, AVRO, HAR file formats and compression.
  • Used Hive to partition and bucket data.
  • Experience in writing MapReduce programs with Java API to cleanse Structured and unstructured data.
  • Wrote Pig Scripts to perform ETL procedures on the data in HDFS.
  • Created HBase tables to store various data formats of data coming from different portfolios.
  • Worked on improving performance of existing Pig and Hive Queries.

Environment: SQL/Server, Oracle 9i, MS-Office, Teradata, Informatica, ER Studio, XML, Business Objects.

Confidential

Data Analyst/Data Modeler

Responsibilities:

  • Deployed GUI pages by using JSP, JSTL, HTML, DHTML, XHTML, CSS, JavaScript, AJAX
  • Configured the project on WebSphere 6.1 application servers
  • Implemented the online application by using Core Java, Jdbc, JSP, Servlets and EJB 1.1, Web Services, SOAP, WSDL
  • Communicated with other Health Care info by using Web Services with the halp of SOAP, WSDL JAX-RPC
  • Used Singleton, factory design pattern, DAO Design Patterns based on the application requirements
  • Used SAX and DOM parsers to parse the raw XML documents
  • Used RAD as Development IDE for web applications.
  • Preparing and executing Unit test cases
  • Used Log4J logging framework to write Log messages with various levels.
  • Involved in fixing bugs and minor enhancements for the front-end modules.
  • Implemented Microsoft Visio and Rational Rose for designing the Use Case Diagrams, Class model, Sequence diagrams, and Activity diagrams for SDLC process of the application
  • Doing functional and technical reviews
  • Maintenance in the testing team for System testing/Integration/UAT
  • Guaranteeing quality in the deliverables.
  • Conducted Design reviews and Technical reviews with other project stakeholders.
  • Was a part of the complete life cycle of the project from the requirements to the production support
  • Created test plan documents for all back-end database modules
  • Implemented the project in Linux environment.

Environment: R 3.0, Erwin 9.5, Tableau 8.0, MDM, QlikView, MLLib, PL/SQL, HDFS, Teradata 14.1, JSON, HADOOP (HDFS), MapReduce, PIG, Spark, R Studio, MAHOUT, JAVA, HIVE, AWS.

Confidential

Data Analyst

Responsibilities:

  • Analyze business information requirements and model class diagrams and/or conceptual domain models.
  • Gather & Review Customer Information Requirements for OLAP and building the data mart.
  • Performed document analysis involving creation of Use Cases and Use Case narrations using Microsoft Visio, in order to present the efficiency of the gatheird requirements.
  • Calculated and analyzed claims data for provider incentive and supplemental benefit analysis using Microsoft Access and Oracle SQL.
  • Analyzed business process workflows and assisted in the development of ETL procedures for mapping data from source to target systems.
  • Worked with BTEQ to submit SQL statements, import and export data, and generate reports in Terra-data.
  • Responsible for defining the key identifiers for each mapping/interface
  • Responsible for defining the functional requirement documents for each source to target interface.
  • Document, clarify, and communicate requests for change requests with the requestor and coordinate with the development and testing team.
  • Coordinated meetings with vendors to define requirements and system interaction agreement documentation between client and vendor system.
  • Generate weekly and monthly asset inventory reports.
  • Coordinate with the business users in providing appropriate, effective and efficient way to design the new reporting needs based on the user with the existing functionality
  • Managed the project requirements, documents and use cases by IBM Rational RequisitePro.
  • Assisted in building an Integrated Logical Data Design, propose physical database design for building the data mart.
  • Document all data mapping and transformation processes in the Functional Design documents based on the business requirements.

Environment: SQL Server 2008R2/2005 Enterprise, SSRS, SSIS, Crystal Reports, Windows Enterprise Server 2000, DTS, SQL Profiler, and Query Analyzer.

We'd love your feedback!