We provide IT Staff Augmentation Services!

Data Scientist Resume

0/5 (Submit Your Rating)

Bloomfield, CT

SUMMARY

  • 6+ years of IT experience in Analysis, Design, Development, Maintenance and Documentation for Data warehouse and related applications using ETL, BI tools, Client/Server and Web applications on UNIX and Windows platforms.
  • Hands on experience with Statistics,DataAnalysis, Machine Learning and deep machine learning using R language.
  • Experienced in creating cutting edgedataprocessing algorithms to meet project demands.
  • Worked with packages like ggplot2 and in R to understanddataand developing applications.
  • Developed predictive models using R to predict customers churn and classification of customers.
  • Worked with heterogeneous relational databases such as Teradata, SQL Server, and MS Access
  • Strong experience in Metadata Management using Hadoop User Experience(HUE)
  • 2 years of experience in Hadoop Metadata Management.
  • Experience in the Data Analysis, Data Mining, Data Mapping, Data Quality, and Data Profiling
  • Expertise in T - SQL in creating and using Views, User Defined Functions, Indexes, Stored procedures involving joins, sub-queries from multiple tables and established relationships between the tables using primary and foreign key constraints
  • Managed different versions of complicated code and distribute them to different teams in the organization utilizing TFS.
  • Developed test packages to simulate errors to prepare for final deployments.
  • Redirected error outputs to error tables to identify improperly populated dimensions and facts.
  • Created complexSSISpackages using various transformations and tasks like Sequence Containers, Script, For loop and ForEach Loop Container, Execute SQL/Package, Send Mail, File System, Conditional Split, Data Conversion, Derived Column, Lookup, Merge Join, Union All, flat file source and destination, OLE DB source and destination, excel source and destination etc.
  • Experienced in performing Incremental Loads and Data cleaning inSSIS. Managed Error handling, logging using Event Handlers inSSIS.
  • Involved in Maintenance and Deployment ofSSISPackages.
  • Developed Custom Reports, Ad-hoc Reports by using SQL Server Reporting Services (SSRS).
  • Monitored Strategies, Processes and Procedures to ensure the Data Integrity, Optimized and reduced Data Redundancy, maintained the required level of security for all production and test databases
  • Create custom data models to accommodate businessmetadataincluding KPIs, Metrics and Goals
  • Create data lineages between business and technicalmetadata, get the lineages reviewed and approved by business users and business system analysts
  • ConductedMetadatatrainings for business users, business analysts and business system analysts
  • Experience on writing, testing and implementing of SQL queries using advance Analytical Functions
  • Knowledge of complete Software Development Life Cycle
  • Administration of the database including performance monitoring, Query tuning & optimization
  • Experience in handling various kind of files (Flat file, CSV and Excel)
  • Experienced in handling concurrent projects and providing expected results in the given timeline
  • Worked on Agile Methodologies and used CA Agile Central

TECHNICAL SKILLS

  • BigData Technologies
  • Hadoop
  • Pig
  • Hive
  • Sqoop
  • HBase
  • Hadoop-Map Reduce
  • Microsoft SQL SERVER 2014/2012/2010
  • Teradata
  • PostgreSQL
  • SSIS/SSAS/SSRS
  • MS Excel
  • MS Visio
  • ERwin
  • SharePoint
  • Windows Server 2012 r2/2008 r2
  • T-SQL
  • Agile and Waterfall

PROFESSIONAL EXPERIENCE

Confidential, Bloomfield, CT

Data Scientist

Responsibilities:

  • Created Domains and Communities for Business Glossary and Information Governance Catalog.
  • Reduced redundancy among incoming incidents by proposing rules to recognize patterns.
  • Worked with Machine learning algorithms like Regressions (linear, logistic etc..), SVMs, Decision trees.
  • Worked on Clustering and classification ofdatausing machine learning algorithms.
  • Estimation and Requirement Analysis of project timelines.
  • Analyzeddataand recommended new strategies for root cause and finding quickest way to solve bigdatasets.
  • Used packages like dplyr, tidyr and ggplot2 in R Studio fordatavisualization.
  • Analyzeddatafrom Primary and secondary sources using statistical techniques to provide daily reports.
  • Developed predictive models using Decision Tree, Random Forest and Naïve Bayes.
  • Created Domains for industry standard code sets and proprietary code sets and maintained reference data by setting up approval workflows.
  • Coordinate withdatascientistsand senior technical staff to identify client's needs and document assumptions.
  • Conducted research on development and designing of sample methodologies, and analyzeddatafor pricing of client's products.
  • Use Correlation analysis to identify relation between variables, patterns, outliers and causal factors.
  • Identified statistically significant variables.
  • Investigated market sizing, competitive analysis and positioning for product feasibility.
  • Worked on Business forecasting, segmentation analysis andDatamining.
  • Used Support vector machines for classification ofdatain groups.
  • Generated graphs and reports using ggplot package in R Studio for analytical models.
  • Worked on the ingestion of data into Hive from different relational databases.
  • Enhanced the tool DMV using Core Java JDBC.
  • Designed data models for Metadata Semantic Layer in ERwin data modeler tool.
  • Reversed Engineered and generated the data models by connecting to their respective databases.
  • Prepared workflows for scheduling the load of data into Hive using IBIS Connections.
  • Automated the process to compare the business metadata with the metahub extract.
  • Designed and developed Looker Reports for the Data Modeling team.
  • Moved reference data, retention data, Information governance catalog, business glossary data to Datalake using Collibra connect.
  • Created a robust automated framework in data lake for metadata management that integrates various metadata sources, consolidates and updates podium with latest and high quality metadata.
  • Ingested data from variety of data sources like Teradata, DB2, Oracle, SQL server and PostgreSQL sources to data lake using podium and solved various data transformation and interpretation issues during the process.
  • Experience in building Data Integration and Workflow Solutions for data warehousing using SQL Server Integration Service (SSIS).
  • Excellent in High Level Design ofETLDTS Packages &SSISPackages for integrating data using OLE DB Connection from heterogeneous sources (Excel, CSV, Oracle, flat file, Text Format Data) by using multiple transformations provided bySSISsuch as Data Conversion, Conditional Split, Bulk Insert, Merge and union all.
  • Responsible for data governance processes and policies solutions using Data Preparation tools and technologies like Podium and Hadoop.
  • Created a data-profiling dashboard in looker by leveraging podium internal architecture, which drastically reduced the time to analyze data quality.
  • Created an analytical model for automating data certification in Data Lake.
  • Created an input agnostic framework for data stewards to handle their ever-emerging work group datasets and created a business glossary by consolidating them.
  • Created a robust comparison process to compare data modelers’ metadata with data stewards’ metadata and identify anomalies.
  • Recommended various technical improvements to teams unfamiliar with big data.
  • Extensive use of GIT as a versioning tool.
  • Worked on Agile Methodologies and used CA Agile Central
  • Trained in R for the statistical analysis in data science and Classification Models which are in used for Machine Learning

Environment: R, R Studio, Podium Data, Data Lake, HDFS, Hue, Hive, Impala, Pig, Looker, ERwin 9.64, HTML, JavaScript, Core Java, PostgreSQL, SSIS, Teradata, R, Classification Models.

Confidential, Livonia, MI

Metadata Analyst/MDM & Data Scientist

Responsibilities:

  • Performeddata analysisanddata profilingusing complexSQLon various sources systems.
  • Created SQL scripts to find data quality issues and to identify keys, data anomalies, and data validation issues
  • Involved in defining the source to target data mappings, business rules, and data definitions
  • Involved in identifying the source data from different systems and map the data into the warehouse
  • Involved in modifying the existing mainframe - Teradata ETL process to Hadoop-Teradata ETL process.
  • Created HBase tables to store variable data formats of input data coming from different portfolios
  • Managed excel spreadsheets, resolved discrepancies associated withmetadata
  • Create technical design documentation for the data models, data flow control process,metadata management.
  • Strong experience in importing themetadatafrom various applications and build end-to-end Data Lineage using ERwin
  • Design and implementmetadatachange management and deployment process while movingmetadatafrom one environment to other
  • Involved in adding huge volumes of data in rows and columns to store data in HBase.
  • Worked with HUE and analyzed the datasets.
  • End-to-end performance tuning of Hadoop clusters and Hadoop Map/Reduce routines against very large data sets.
  • Successfully loaded files to Hive and HDFS from Teradata Database.
  • Importing and exporting data into HDFS and Hive from Teradata using Sqoop.
  • Involved in creating Hive tables, loading with data and writing hive queries, which will run internally in map, reduce way.
  • Experienced in managing andreviewingHadooplog files.
  • Load and transform large sets of structured data.
  • Responsible to manage data coming from different sources.

Environment: BigData Technologies (Hadoop, Pig, Hive, Sqoop, HBase, Hadoop-Map reduce)MS SQL Server 2014/2008, SSIS Package, SQL BI Suite (SSMS, SSIS, SSRS, SSAS), XML, MS Excel, MS Access 2013Windows Server 2012, SQL Profiler, Erwin r 7.3., Net 4.5, TFS 2013.

Confidential, Southfield, MI

Data Analyst/SQL Developer

Responsibilities:

  • Worked on EDW (Enterprise Data Warehouse) for Claims Reporting Project.
  • Analyzed the datasets and loaded the data into SQL Server tables for BI reporting project
  • Ensured best practices are applied and integrity of data is maintained through security and documentation
  • Configured and Maintained Report Manager and Report Server for SSRS.
  • Created reports to retrieve data using Stored Procedures that accept parameters depending upon the client requirements
  • Involved in Debugging and Deploying reports on the production server
  • Involved in data management processes and ad hoc user requests
  • Wrote standard T-SQL to perform Data Validation and create Excel summary reports (pivot tables and charts).
  • Involved in requirements gathering, source data analysis and identified business rules for data migration and for developing data warehouse/data mart
  • Involved in identifying and defining the data inputs and captured metadata and associated rules from various source of data for ETL Process for data warehouse
  • Worked with Business Analyst to develop business rules that support the transformation of data
  • Involved the verification of data accuracy within SAS Analytical Systems and source systems
  • Worked with SAS Datasets and analyzed the data for analytical reporting
  • Created Source to target mapping documents from staging area to Data Warehouse
  • Designed and optimized indexes, views, stored procedures and functions using T-SQL
  • Helped designing and implementing processes for deploying, upgrading, managing, archiving and extracting data for reporting
  • Performed maintenance duties like performance tuning and optimization of queries, functions and stored procedures

Environment: SAS 9.2, SAS Datasets, SQL Server 2008/2012, SSIS, SSRS, Tidal Job Scheduler

Confidential

ETL Developer

Responsibilities:

  • Participated in the Software Development Life Cycle (SDLC) processes including Analysis, Design, Coding, Testing and Deployment.
  • Created Cubes with Dimensions and Facts and calculating measures and dimension members using Multidimensional expression (MDX).
  • Involved in building the Enterprise Data Warehouse for Policy Management, Premium Billing, Member Policy Renewal process.
  • Created SQL objects like Tables, Stored Procedures, Functions, and User Defined Data-Types.
  • Created SSIS packages to load data into Data Warehouse using Various SSIS Tasks like Execute SQL Task, bulk insert task, data flow task, file system task.
  • Developed ETL jobs to load information into Data Warehouse from different relational databases and flat files.
  • Generated different reports using MS SQL Reporting Services
  • Created various Tabular and Matrix Reports using SSRS.

Environment: SQL Server 2005/2008, SSIS, SSRS, SQL Stored Procedures, Flat files, Excel, MS Office

We'd love your feedback!