Data Scientist Resume
Dallas, TX
SUMMARY
- Data Scientist wif 8+ years’ experience in Data mining, Statistical Modeling, Data Science, Data Acquisition, Data Provisioning, Predictive Analytics, Data Visualization, and Machine Learning.
- Experienced in transforming business requirements into analytical models to serve business needs.
- Experience working wif Machine Learning techniques for supervised learning models such as linear regression, logistic regression, decision trees, random forests, linear discriminant analysis
- Experience in Data preparations by exploration, cleansing, imputing missing values and normalization
- Experience in Exploratory Data Analysis by conducting variable identification for variance differences and performing cluster analysis to identify homogenous groups wifin teh data
- Good knowledge of statistical analysis techniques such as confidence interval, hypothesis testing, ANOVA, correlation, regression, time series,
- Experience in conducting dimensionality reduction techniques like TEMPprincipal component analysis(PCA) to assess collinearity, Linear discriminant analysis(LDA) and independent component analysis.
- Gained knowledge in neural networks like CNN, RNN, DNN wif understanding of RNN and LSTM using tensor flow and Keras
- Experience in developing data mining and reporting solutions that scale across massive volumes of data
- Gained working knowledge of natural language processing techniques like word2vec, BOW (Bag of words), Avg word2Vev, weighted word2Vec problems
- Experience wif Big Data Architecture in data provisioning like Hadoop, HDFS, Hive, Hue along wif computing engines like Spark and interactions wif nginx server for application and Data transactions
- Working knowledge of Computing servers like Spark, Photon, AWS - EMR, SAS SPRE, CAS Server, Hive.
- Experience working wif file types such as avro, parquet, gzip, snappy zip files for optimized compute resource consumption and aggregate collection wifin teh data lake and Data warehouse.
- Experience wif AWS cloud architecture such as S3 for storage, EC2, EMR, EBS, Redshift, RDS and Visualization wif Quicksight, Tableau, PowerBI, Qlikview and in-code visualization wif libraries plotly, seaborn, matplotlib
- Proficient wif relational Databases like MySQL and noSQL databases like MongoDB.
- Proficient wif Data Visualization tools such as Tableau, Power BI and Quicksight.
- Experience in IDE like Anaconda, Jupyter Notebook, Visual Studio code, IntelliJ, PyCharm and Spyder.
- Experience in code management solutions like Git, Github, Git Bash.
- Experience in wrangling Data using Python Libraries on Jupyter Notebook such as Pandas, Numpy, Matplotlib, scikit learn, seaborn, pyspark, plotly for manipulation and Visualization.
- Experience wif SAS Viya architecture in big data provisioning, Cleaning, Exploratory Data Analysis and Principle component analysis on Cloud analytics services and visualization using Visual Analytics
- Supervised Source to target mappings and data validation for data ingest operations.
- Proficient in writing complex SQL queries wif joins and subqueries to access and manipulate database systems like MySQL, PostgreSQL, PL/SQL and snowflake.
- Proficient in entire project lifecycle phases including data acquisition, provisioning, engineering, feature engineering, statistical modeling, Data Visualization and Machine Learning Modeling
TECHNICAL SKILLS
Database Mgmt: MySQL, PostgreSQL, Hive, RedShift, RDS, DAG, Data dictionary, STTM BI, Visualization, Data Analysis, Data Wrangling, Machine Learning, reporting tools Tableau, Qlik view, Qlik sense, Crystal reports, Power BI, AWS Quicksight. Trifacta, Arcadia Data, Talend, Tibco Spotfire, Power BI, Tableau, Hue, SAS Viya, Minitab SQL, Python (NumPy, pandas, PySpark, scikit-learn, statsmodels, matplotlib, plotly, seaborn) Linear regression, Logistic regression, Classification Methods, ensemble methods, Regular expression, decision trees, random forests, clustering algorithms, Naïve Bayes, Support vector machines, K- means Clustering, Probability clustering, Neural networks, TEMPprincipal component analysis, independent component analysis, Natural Language processing, Interactive voice response, word processors
Big Data, Cloud Computing, Distributed Computing Tools: Apache Hadoop, HDFS, Hive, Avro, Parquet, nginx, SparkAWS: S3, EMR, EC2, Quicksight, Redshift, RDS, GlueSAS: Cloud Analytics server CAS, Viya, Visual Analytics, Lineage viewer, Data Studio, Model Studio, Graph Builder, Environment Manager, Business Objects, Data Warehousing, Net Weaver
DB Environment: MS Access, SQL server, MS SQL server studio, Hive, AWS RDS, AWS Redshift, Hadoop, Spark, EMR
Project Mngt.: MS Project, Primavera, MS Office Suite, MS Visio
Prog. Languages: SQL, Python, R, SAS, SPSS, C, Java, HTML, CSS, JS, XML
MS Office Suite: Excel, Access, Outlook, Word, Power point, Skype, Lync, OneNote
PROFESSIONAL EXPERIENCE
Confidential, Dallas, TX
Data Scientist
Responsibilities:
- Involved in project planning and collecting historical data about most frequent commands used by users.
- Lead activities while interacting wif various cross functional teams across data engineering, software engineering in gathering requirements and developing a roadmap for predictive analytics
- Supported Updating existing models by retraining them to support current data and infrastructures.
- Gatheird business needs from interactions wif cross functional teams to build data driven strategies
- Conducted Data Preparation, descriptive data analysis, feature engineering, variable selection and Exploratory Data Analysis to identify teh key trends on call intent prediction model using logistic regression, decision tree, random forests and deep learning algorithms like neural networks.
- Utilized regular expressions to clean teh existing user logs to limit teh vocabulary generation
- Performed text analysis on teh data using natural language processing techniques like bag of words, term frequency inverse document frequency, Word2vec, average word2vec wif teh help of NLTK libraries
- Supported building multi-linear classification algorithm for real-time text mining on extracted data
- Created user personas to extract information, customer behavior & new attributes for anomaly detection
- Developed insights due to correlation between teh customer journey and NLP model outputs.
- Implemented existing vocabulary commands from past data for building user word dictionary
- Implemented cosine similarity to pick directional similarities in teh patterns of commands as a measure
- Worked wif Data engineering and cloud infrastructure team to deploy infrastructure for teh data strategy developed towards teh business requirements and meet business goals
- Gained experience on natural language processing to understand teh existing sentiments of users such that teh outcome is applied to their existing services
- Built visualization on Tableau and interactive dashboards to present teh results for team members
- Geo and demographic segmentation of users to efficiently allocate resources in future expansion.
Environment: Hadoop, Hive, Spark, Python, Anaconda, Jupyter, Logistic Regression, Random forest, Gradient boosting, sampling techniques, transformations, Shell, Hue, Python, Numpy, Pandas, scikit-learn, plotly, matplotlib, seaborn.
Confidential, Duluth, GA
Data Engineer/ Data Scientist
Responsibilities:
- Sourced historical data from relevant and novel data sources hosted on SQL databases to understand and develop patterns for customers’ purchases from time to time.
- Analyzed purchase data for customers and trending products to build a recommendation for teh types of products and services to customers based on their behavior tracked through teh customer accounts, purchase history and geo location.
- Optimized product deliveries by avoiding various considerations related to product inventory floor stocks
- Involved in developing business strategy wif deals and discount coupons to offer improved customer purchases by employing pattern recognition algorithms
- Segmented customer groups to further analyze behavioral patterns of customers using K-means clustering and hierarchical clustering methods on customer actions
- Employed multi-linear regression to generate customer lifetime value from data curated quarterly
- Developed model to collect data to optimize product placement in order to get attention from shoppers
- Data is segregated to fit teh right customer profile and offer high demand products and reduce clutter
- Involved in running deep learning models for teh process of comparing items against each other, tracking their performance and suggest business groups to support decisions
- Involved in building one class Support Vector Machine (SVM) and TEMPPrincipal Component Analysis (PCA) algorithms for anomaly detection of fraud and other errors that signal dishonest behaviors.
- Forecasted sales and improved accuracy - (MILP) by 30% by implementing advanced forecasting algorithms that were effective in detecting seasonality and trends in teh patterns.
- Supported in developing demand forecast algorithm for new fashion products and generate initial demand forecast based on teh historical product and store wif periodical updates based on sales data
- Utilized random forest regression model to forecast product sales distribution among selected centers.
- Analyzed price cuts and discounted products using price elasticity of demand for products wif regression
- Suggested selective and flagged price cuts for certain licensing categories that fell under threshold.
- Machine learning algorithms were implemented using python Sci-kit learn, SciPy, NumPy, Pandas, seaborn modules to analyze teh terabytes of data.
- Published visuals and reports using QlikView and QlikSense objects like tables, Pivot tables, Containers, Line charts, Bar charts, Combo charts, Scatter charts, Line objects, Pie charts utilized by business users.
- Created PDF push reports and setup email distribution on QlikView Publisher 9
Environment: Python, Jupyter notebook, Spark, Hadoop, Hive, SQL, Hue, Libraries like pandas, matplotlib, plotly, numpy, seaboarn, skikit-learn
Confidential, Newark, NJ
Stat Modeler / Data Scientist
Responsibilities:
- Responsible for data aggregation, pre-processing, cleansing, imputing missing values and descriptive statistics
- Worked closely wif subject matter experts and business analysts to draw metrics for inferential statistics for predictive analytics and prescriptive patterns using data to build business solutions
- Analyzed teh data using various machine learning algorithms to develop a pay model using customer DNA and past payment history to offer better deals to teh customers before billing cycles
- Gatheird historical data from sales and other heterogeneous data sources to analyze customer patterns
- Performed feature engineering by feature selection and replacing missing values and data normalization
- Different models of conditional random fields are experimented wif metrics such as confusion matrix, F scores and ROC curves
- To maximize teh business overall benefit to cost, teh model is trained to identify teh customers who need it, which require decreasing false positives while having enough accuracy for teh model.
- Identified customer’s ability-to-pay through customer calls and reevaluate and improve teh business services to teh customer
- Based on teh ability-to-pay, analyzed teh causes for inability to pay through customer calls
- Extensive text cleansing is done by going through teh stop words, word replacement and lemmatization
- Identified teh keywords using latent semantic analysis and performed truncated semantic analysis that are related to teh inability to pay model
- Classification of keywords in teh transcripts are done based on their appearances and used as another feature as previous call reason in teh ability to pay model.
- Insights from teh models are visualized on tableau and shared wif teh leadership and business users
- Insights from teh model are used for customizing teh business services provided by Confidential .
Environment: Data Aggregation, DNA, Transcripts, Tableau, historical data, Confidential .
Confidential, Tampa, FL
Senior Data Analyst
Responsibilities:
- Utilized Data source to target mapping resources to develop source-to-target mapping documents
- Involved in Data mapping specifications based on data dictionary and data catalogue
- Produced business requirement documents based on interviews wif business users and technical users
- Analyzed system requirements, Data requirements, Data dictionaries, Entity relationship diagrams, data mapping requirement specifications and produced functional requirements for teh data pipeline
- Produced Data Quality metrics for teh Data center using functional requirements and supplementary documents for both technical and non-technical team members
- Analyzed various ETL mappings and sessions as per business requirements and business rules to load required data from source systems to target systems
- Created a test environment for staging area, loading teh staging area wif data from multiple sources
- Created test cases for validating data sourced from source systems and making way into target systems
- Compiled Data Validation criteria for teh Data Quality test cases that are used to test data ingestion
- Analyzed various data sources wif flat files, ASCII Data, Relational Data (Oracle, DB2 UDB, MS SQL Server) and other heterogeneous data sources
- Successfully tested several advanced SQL processes such as stored procedures and triggers for ETL using cases, having, connect by, exec, declare, instead of, for, after etc
- Developed Unit testing and Performance tuning to ensure testing issues were resolved on teh basis of defect reports for Teradata SQL development
- Validated data in several scenarios such as before and after teh ETL process to ensure ETL scripts quality
- Performed testing on messages that are published by ETL tool and data being loaded into databases
- Gained experience in creating UNIX scripts for file transfer and file manipulation
- Coordinated between onsite and offsite teams for data transition and managed file transfers by conducted quality assurance processes and validating teh data for use.
- Validated database transactions for fields, lengths, file sizes, constraints, stored procedures, cross walks and datatypes defined in teh data catalogue and metadata information resources.
Environment: Informatica 7.1, Data Flux, Oracle 9i, Quality Center 8.2, SQL, TOAD, PL/SQL, FLatfiles
Confidential
Data Analyst
Responsibilities:
- Analyzed incoming data by performing descriptive statistics and explored various elements for integrity wifin in teh data and TEMPhas potential to be used in analysis
- Monitored and resolved data flow issues on a routine basis by conducting data quality checks
- Explored data using graphical representations such as histograms, box plots, skewness, and outliers
- Extensively created excel analysis for business rules and decisions for forecasting
- Visualized data through Excel charts, pivot tables and functions
- Experienced working wif SQL server and created tables and views along wif using queries such as WHERE, HAVING, JOINS, functions and stored procedures
- Created views for teh marketing team for reporting marketing numbers by using marketing data
- Worked wif cross functional teams to resolve data discrepancies and logical data corrections in teh reports
- Implemented metadata models for reporting functionalities and automated correction process
- Performed ETL processes to extract transform and load data from OLTP system to OLAP system using Talend
- Reviewed logical model wif application developers, ETL team, DBA team and testing team to provide information about data model and business requirements
- Acquired data from primary and secondary data sources to conduct analysis
- Created filters, parameters and calculated sets for preparing dashboards and worksheets in tableau
- Generated ad hoc reports using excel sheets, flat files and CSV files
- Worked wif management team to create prioritized list of needs for each business segment
- Facilitated new client to transition into new data platform and migrated auxiliary resources
Environment: Histograms, Box Plots, Skewness, Excel analysis, Excel charts, pivot table, OLTP, ETL, CSV files.
