We provide IT Staff Augmentation Services!

Data Scientist Resume

3.00/5 (Submit Your Rating)

Chicago, IL

PROFESSIONAL SUMMARY:

  • Over 8+ years of experience in areas including Data Analyst, Statistical Analysis, Machine Learning, Data mining with large data sets of structured and unstructured data in banking, travel services, strong functional knowledge, business processes and latest market trends and manufactory industries.
  • Proficient in Predictive Modeling, Data Mining Methods, Factor Analysis, ANOVA, Hypothetical testing, normal distribution and other advanced statistical and econometric techniques.
  • Developed predictive models using Decision Tree, Random Forest, Naïve Bayes, Logistic Regression, Cluster Analysis, and Neural Networks.
  • Experienced the full software lifecycle in SDLC, Agile, and Scrum methodologies.
  • Strong SQL programming skills, with experience in working with functions, packages, and triggers.
  • Excellent understanding of machine learning techniques and algorithms, such as k - NN, Naive Bayes, SVM, Decision Forests, natural language processing (NLP) etc.
  • Worked with RDBMS including MySQL, DB2 and Oracle SQL.
  • Experienced in Data Integration Validation and Data Quality controls for ETL process and Data Warehousing using MS Visual Studio SSIS, SSAS, SSRS.
  • Experience in implementation of the Stored Procedures, Triggers, Functions using T-SQL
  • Expert in developing Data Conversions/Migration from Legacy System of various sources (flat files, Oracle, Non-Oracle Database) to Oracle system Using SQL Loader, External table and Calling Appropriate Interface tables and API's Informatica.
  • Experienced in Performance tuning of Informatica (sources, mappings, targets, and sessions)
  • Hands on experience in Teradata SQL Analytics, Teradata utilities and familiar in Creating Secondary indexes, and join indexes in Teradata.
  • Strong working experience on Teradata query performance tuning by analyzing CPU, AMP Distribution, Table Skewness, and IO metrics.
  • Excellent Data Mining skills who can sift the grain from large datasets of Structured and Unstructured data, identify the patterns within data, analyze data and interpret results into actionable insights and business values
  • Well adapted to Statistical Programming Languages and adept at writing code in R, Python, SAS and cloud platform as Azure ML and AWS ML
  • Experience in managing and maintaining IAM policies for organizations in AWS to define groups, create users, assign roles and define rules for role-based access to AWS resources.
  • Hands on experience in setting up databases in AWS using RDS, storage using an S3 bucket and configuring instance backups to S3 bucket to ensure fault tolerance and high availability.
  • Maintenance and monitoring of Docker in a cloud-based service during production and Set up a system for dynamically adding and removing web services from a server using Docker. Used Kubernetes to manage Docker containers cluster.
  • Configuration management using Amazon Cloud Formation, Continuous integration with Jenkins, AWS management (EC2, EBS, RDS, Route53)
  • Transformed traditional environment to virtualized environments with AWS-EC2, S3, EBS, ELB, and EBS.
  • Skills to build a fully automated, highly elastic cloud orchestration framework on AWS.
  • Extensively worked on Teradata Utility tools like BTEQ, Fast load, Fast Export, Multi-Load, TPUMP, and TPT.
  • Proficient in Tableau and R-Shiny data visualization tools to analyze and obtain insights into large datasets create visually powerful and actionable interactive reports and dashboards.
  • Automated recurring reports using SQL and Python and visualized them on BI platform like Tableau.
  • Worked in a developmentenvironment like Git and VM.
  • Experience in developing and analyzing data models, involved in writing simple and complex SQL queries to extract data from the database for data analysis and testing
  • Strong knowledge in all phases of the SDLC (Software Development Life Cycle) from analysis, design, development, testing, implementation, and maintenance with timely delivery against deadlines
  • Ability to thoroughly analyze the system's functional requirements and prepare BRD (Business requirement documentation), use cases and testing documents
  • Expertise in defining thescope of the project post gathering business requirements including constraints, assumptions, business impacts, project risks & scope exclusions, conducting a GAP analysis, User Acceptance Testing (UAT, SWOT Analysis, Cost-benefit analysis and ROI analysis)
  • Proficient knowledge of statistics, mathematics, machine learning, recommendation algorithms and analytics with an excellent understanding of business operations and analytics tools for effective analysis of data.
  • Excellent communication skills. Successfully working in fast-paced multitasking environment both independently and in a collaborative team, a self-motivated enthusiastic learner.

TECHNICAL SKILLS:

Programming & Scripting Languages: C, C++,PL/SQL.

MS: Access, Oracle 12c/11g/10g/9i,Mysql,DB2

Statistical Software: SPSS, R, SAS.

ETL/BI Tools: Informatica Power Center 9.x, Tableau, Cognos BI 10, MS Excel, SAS, SAS/Macro, SAS/SQL

Cloud: AWS, S3, EC2.

Statistical Methods: Time Series, regression models, splines, confidence intervals, principal component analysis and Dimensionality Reduction, bootstrapping

BI Tools: Microsoft Power BI, Tableau, SSIS, SSRS, SSAS, Business Intelligence Development Studio (BIDS), Visual Studio, Crystal Reports, Informatica 6.1.

Data Modeling: Kimball/Inmon, Logical/Physical/Dimensional, Star/Snowflake Schema, ETL, OLAP, Waterfall, Agile.

Data Modelling: Erwin r 9.6, 9.5, 9.1, 8.x, Rational Rose, ER/Studio, MS Visio, SAP Power designer.

BTEQ, Fast load, Fast Export, Multi: load, TPUMP and TPT

Database Tools: Toad, SQL Developer, PL/SQL Developer, SQL Developer, SQL*Loader, SQL*Plus, Informatica Power Center 9.5.1.

Operating Systems: Windows (10, 7, Vista), XP, UNIX, Linux.

PROFESSIONAL EXPERIENCE:

Confidential, Chicago, IL

Data Scientist

Responsibilities:

  • Worked as a Data Modeler/Analyst to generate Data Models using Erwin and developed a relational database system.
  • Analyzed the business requirements of the project by studying the Business Requirement Specification document.
  • Extensively worked on Data Modeling tools Erwin Data Modeler to design the data models.
  • Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
  • A highly immersive Data Science program involving Data Manipulation & Visualization, Web Scraping, Machine Learning, SQL, GIT, Unix Commands, NoSQL, MongoDB.
  • Transformed Logical Data Model to Erwin, Physical Data Model ensuring the Primary Key and Foreign Key relationships in PDM, Consistency of definitions of Data Attributes and Primary Index Considerations.
  • Designed Mapping to process the incremental changes that exist in the source table. Whenever source data elements were missing in source tables, these were modified/added inconsistency with third normal form based OLTP source database.
  • Performed social network analysis and topic modeling in R, on employee chat data, and develop Sankey plot to understand the communication paths, the strength of relations between Agents and the topics frequently discussed between them.
  • Analyzed employee behavior and performance data, and developed Shiny dashboards to evaluate team preparedness through metrics, which helped evaluate leadership skills, agent experiences, agent behavior and customer sentiments
  • Developed SQL procedures to synchronize the dynamic data generated from GTID systems with the Azure SQL Server.
  • Extensively use Python's multiple data science packages like Pandas, NumPy, matplotlib, Seaborn, SciPy, Scikit-learn, and NLTK.
  • Work on data pre-processing and cleaning the data to perform feature engineering and performed data imputation techniques for the missing values in the dataset using Python.
  • Implement machine learning model (logistic regression, XGBoost, SVM) with Python Scikit- learn.
  • Work on different data formats such as JSON, XML and applied machine learning algorithms in Python.
  • Performed Exploratory Data Analysis, trying to find trends and clusters.
  • Develop rigorous data science models to aggregate inconsistent real-time signals into strong predictors of market trends.
  • Constantly monitor the data and models to identify the scope of improvement in the processing and business. Manipulated and prepared the data for data visualization and report generation. Performed data analysis, statistical analysis, generated reports, listings, and graphs.
  • Designed tables and implemented the naming conventions for Logical and Physical Data Models in Erwin 7.0.
  • Provide expertise and recommendations for physical database design, architecture, testing, performance tuning and implementation.
  • Designed logical and physical data models for multiple OLTP and Analytic applications.
  • Extensively used the Erwin design tool &Erwin model manager to create and maintain the Data Mart.
  • Designed the physical model for implementing the model into the oracle9i physical database.
  • Involved with Data Analysis Primarily Identifying Datasets, Source Data, Source Meta Data, Data Definitions and Data Formats
  • Performance tuning of the database, which includes indexes, and optimizing SQL statements, monitoring the server.
  • Wrote simple and advanced SQL queries and scripts to create standard and Adhoc reports for senior managers.
  • Collaborated the data mapping document from a source to target and the data quality assessments for the source data.
  • Created S3 buckets and managed roles and policies for S3 buckets. Utilized S3 buckets and Glacier for file storage and backup on AWS cloud. Used Dynamo DB to store the data for metrics and backend reports.
  • Worked with Elastic Beanstalk for quick deployment of services such as EC2 instances, Load balancer, and databases on the RDS on the AWS cloud environment.
  • Used Amazon Simple Workflow service (SWF) for data migration in data centers which automates the process and tracks every step and logs are maintained in S3 bucket.
  • Designed and developed user interfaces and customization of Reports using Tableau and OBIEE and designed cubes for data visualization, mobile/web presentation with parameterization and cascading.
  • Performed Data Analysis and Data Profiling and worked on data transformations and data quality rules.
  • Created SSIS Packages using Pivot Transformation, Execute SQL Task, Data Flow Task, etc. to import data into the data warehouse.
  • Developed and implemented SSIS, SSRS and SSAS application solutions for various business units across the organization.

Environment: SQL, GIT, Unix Commands, NoSQL, MongoDB, SSIS, SSRS, SSAS, AWS,S3,EC2,RDS,SWF,Dynamo DB, Glacier, Erwin, Tableau, OBIEE.

Confidential, Michigan

Data Scientist

Responsibilities:

  • Data collection, Cleaned, filtered and transformed data to the specified format.
  • Created captivating interactive visualizations and presentations to enhance decision-making capabilities of the management.
  • Developed novel applications of classification, forecasting, simulation, optimization, and summarization techniques to enhance effective decisions.
  • Prepared the workspace for Markdown. Accomplished Data analysis, statistical analysis, generate dreports, listings, and graphs.
  • Found outliers, anomalies, trends in any given data sets.
  • Assisted in migrating data, data pump with the Export/Import utility tool.
  • Provided daily change management process support, ensuring that all changes to program baselines are properly documented and approved, maintained, managed and issue change schedules.
  • Developed, installed, maintained and monitored company databases in high performance/high availability environment with supported configuration, performance tuning to ensure optimal resource usage.
  • Documented all programs and procedures to ensure an accurate historical record of work completed on the assigned project as well as to improve quality and efficacy.
  • Implemented various types of change data captures according to source data behavior and business requirements.
  • Implemented various Performance tuning techniques at ETL & Teradata BTEQ for efficient development and performance.
  • Used Simple storage services (s3) for snapshot and Configured S3 lifecycle of Applications & Databases logs, including deleting old logs, archiving logs based on retention policy of Apps and Databases.
  • Created and maintained Logical and Physical models for the data mart and created partitions and indexes for the tables in the data mart.
  • Performed Data profiling and Analysis applied various data cleansing rules designed data standards and architecture/designed the relational models.
  • Creating new data designs and ensuring that they fall within the realm of the overall Enterprise BI Architecture.
  • Built models using Statistical techniques like Bayesian HMM and Machine Learning classification models like XG Boost, SVM, and Random Forest.
  • Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
  • Created logical data model from the conceptual model and its conversion into the physical database design using Erwin 9.6.
  • Designed and developed new reports and maintained existing reports using Microsoft SQL Reporting Services (SSRS) and Microsoft Excel to support the firm's strategy and management.
  • Created sub-reports, drill down reports, summary reports, parameterized reports, and ad-hoc reports using SSRS.
  • Used SAS/SQL to pull data out from databases and aggregate to provide detailed reporting based on the user requirements.
  • Used SAS for pre-processing data, SQL queries, data analysis, generating reports, graphics, and statistical analyses.
  • Use SAS statistical regression method and SAS/REG polynomial simulation in Excel to simulate the anisotropic trend as 1D depth functions. Validate the simulated function by the image quality of depth migration.
  • Tested the migrated data processing system on Google Cloud with velocity model updating tasks.
  • Designed and Developed Oracle PL/SQL and Shell Scripts, Data Import/Export, Data Conversions and Data Cleansing.
  • Responsible for the development of target data architecture, design principles, quality control, and data standards for the organization.
  • Worked with DBA to create Best-Fit Physical Data Model from the Logical Data Model using Forward Engineering in Erwin.
  • Produced quality reports for management for decision making.
  • Participated in all phases of research including data collection, data cleaning, data mining, developing models and visualizations.
  • Redefined many attributes and relationships and cleansed unwanted tables/columns using SQL queries.
  • Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.

Environment: ETL, Teradata BTEQ, S3, XGBOOST, SVM, Random Forest, AWS, Oracle PL/SQL, Erwin 9.6, DBA, SQL, Shell Script, HMM, Spark SQL, PySpark.

Confidential, Houston, TX

Jr Data Scientist

Responsibilities:

  • Developed scalable machine learning solutions within a distributed computation framework (e.g. Hadoop, Spark, Storm etc.).
  • Utilizing NLP applications such as topic models and sentiment analysis to identify trends and patterns within massive data sets.
  • Creating automated anomaly detection systems and constant tracking of its performance
  • Strong command of data architecture and data modeling techniques.
  • Knowledge in ML & Statistical libraries (e.g. Scikit-learn, Pandas).
  • Having the knowledge to build predictive models to forecast risks for product launches and operations and help predict workflow and capacity requirements for TRMS operations
  • Having experience with visualization technologies such as Tableau
  • Draw inferences and conclusions, and create dashboards and visualizations of processed data, identify trends, anomalies
  • Generation of TLFs and summary reports, etc. ensuring on-time quality delivery.
  • Participated in client meetings, teleconferences and video conferences to keep track of project requirements, commitments made and the delivery thereof.
  • Solved analytical problems, and effectively communicate methodologies and results
  • Worked closely with internal stakeholders such as business teams, product managers, engineering teams, and partner teams.
  • Data mining using state-of-the-art methods
  • Extending the company's data with third-party sources of information when needed
  • Enhancing data collection procedures to include information that is relevant for building analytic systems
  • Processing, cleansing, and verifying the integrity of data used for analysis
  • Doing the ad-hoc analysis and presenting results in a clear manner.
  • Created automated metrics using complex databases.
  • Foster culture of continuous engineering improvement through mentoring, feedback, and metrics.

Environment: Erwin r9.0, Informatica 9.0, ODS, OLTP, Oracle 10g, OLAP, DB2, Metadata, MS Excel, Mainframes MS Visio, Rational Rose, PL/SQL, etc.

Confidential, Santa Clara, CA

Data Analyst/Modeler

Responsibilities:

  • Developed the logical data models and physical data models that capture current state/future state data elements and data flows using ER Studio.
  • Delivered dimensional data models using ER/Studio to bring in the Employee and Facilities domain data into the Oracle data warehouse.
  • Developed the design & Process flow to ensure that the process is repeatable.
  • Performed analysis of the existing source systems (Transaction database)
  • Involved in maintaining and updating the Metadata Repository with details on the nature and use of applications/data transformations to facilitate impact analysis.
  • Created DDL scripts using ER Studio and source to target mappings to bring the data from source to the warehouse.
  • Designed the ER diagrams, logical model (relationship, cardinality, attributes, and, candidate keys) and physical database (capacity planning, object creation, and aggregation strategies) for Oracle and Teradata.
  • Worked in importing and cleansing of data from various sources like Teradata, Oracle, flat files, MS SQL Server with high volume data
  • Designed Logical & Physical Data Model /Metadata/ data dictionary using Erwin for both OLTP and OLAP based systems.
  • Reverse Engineered DB2 databases and then forward engineered them to Teradata using ER Studio.
  • Part of a team conducting logical data analysis and data modeling JAD sessions communicated data-related standards
  • Involved in meetings with SME (subject matter experts) for analyzing the multiple sources.
  • Created DDL scripts using ER Studio and source to target mappings to bring the data from source to the warehouse.
  • Identify, assess and intimate potential risks associated with testing scope, quality of the product and schedule
  • Wrote and executed SQL queries to verify that data has been moved from a transactional system to DSS, Data warehouse, data mart reporting system in accordance with requirements.
  • Worked in importing and cleansing of data from various sources like Teradata, Oracle, flat files, SQL Server 2005 with high volume data.
  • Wrote and executed SQL queries to verify that data has been moved from a transactional system to DSS, Data warehouse, data mart reporting system in accordance with requirements.
  • Worked in importing and cleansing of data from various sources like Teradata, Oracle, flat files, SQL Server 2005 with high volume data
  • Worked extensively on ER Studio for multiple Operations across Atlas Copco in both OLAP and OLTP applications.
  • Generated comprehensive analytical reports by running SQL queries against current databases to conduct data analysis.
  • Co-ordinate all teams to centralize Meta-data management updates and follow the standard Naming Standards and Attributes Standards for DATA &ETL Jobs.
  • Finalize the naming Standards for Data Elements and ETL Jobs and create a Data Dictionary for Meta Data Management.

Environment: ER Studio, Business Objects XI, Rational Rose, Data stage, MS Office, MS Visio, SQL, SQL Server 2000/2005, Rational Rose, Crystal Reports 9, SQL Server 2008, SQL Server Analysis Services, SSIS, Oracle 11g.

Confidential

Data Analyst

Responsibilities:

  • Used SAS Proc SQL pass-throughfacility to connect to Oracle tables and created SAS datasets using various SQL joins such as left join, right join, inner join and full join.
  • Performing data validation, transforming data from RDBMS oracle to SAS datasets.
  • Produce quality customized reports by using PROC TABULATE, PROC REPORT Styles, and ODS RTF and provide descriptive statistics using PROC MEANS, PROC FREQ , and PROC UNIVARIATE.
  • Developed SAS macros for data cleaning, reporting and to support routing processing.
  • Performed advanced querying using SAS Enterprise Guide, calculating computed columns, using a filter, manipulate and prepare data for Reporting, Graphing, and Summarization, statistical analysis, finally generating SAS datasets.
  • Involved in Developing, Debugging, and validating the project-specific SAS programs to generate derived SAS datasets , summary tables, and data listings according to study documents.
  • Created datasets as per the approved specification collaborated with project teams to complete scientific reports and review reports to ensure accuracy and clarity.
  • Experienced in working with data modellers to translate business rules/requirements into conceptual/logical dimensional models and worked with complex de-normalized and normalized data models
  • Created action filters, user filters, parameters and calculated sets for preparing dashboards and worksheets in Tableau.
  • Combined Tableau visualizations into Interactive Dashboards using filter actions, highlight actions etc. and published them on the web.
  • Gathering business requirements, creating business requirement documents ( BRD /FRD ).
  • Work closely with business leaders and users to define and design the data sources requirements and data access Code, test, identify, implement and document technical solutions utilizing JavaScript, PHP & MySQL.
  • Created Rich dashboards using Tableau Dashboard and prepared user stories to create compelling dashboards to deliver actionable insights
  • Working with the manager to prioritize requirements and preparing reports on the weekly and monthly basis.

Environment: SQL Server, Oracle 11g/10g, MS Office Suite, Power Pivot, Power Point, SAS Base, SAS Enterprise Guide, SAS/MACRO, SAS/SQL, SAS/ODS, SQL, PL/SQL, Visio

Confidential

Data Analyst

Responsibilities:

  • Implemented Microsoft Visio and Rational Rose for designing the Use Case Diagrams, Class model, Sequence diagrams, and Activity diagrams for SDLC process of the application
  • Worked with other teams to analyze customers to analyze parameters of marketing.
  • Conducted Design reviews and Technical reviews with other project stakeholders.
  • Was a part of the complete life cycle of the project from the requirements to the production support.
  • Created test plan documents for all back-end database modules
  • Used MS Excel, MS Access, and SQL to write and run various queries.
  • Used a traceability matrix to trace the requirements of the organization.
  • Recommended structural changes and enhancements to systems and databases.
  • Conducted Design reviews and Technical reviews with other project stakeholders.
  • Maintenance in the testing team for System testing/Integration/UAT
  • Guaranteeing quality in the deliverables.

Environment: UNIX, SQL, Oracle 10g, MS Office, MS Visio.

We'd love your feedback!