We provide IT Staff Augmentation Services!

Data Analyst/ Data Scientist Resume

0/5 (Submit Your Rating)

Fairfax, VA

SUMMARY

  • Around 7+ years of experience in inData Analytics, experience in data scrubbing to mine the data, Data Acquisition, Data Engineering to extract features utilizing Statistical Techniques, Exploratory Data Analysis with an inquisitive mind, build diverseMachine Learning Algorithms for developing Predictive Models and design Stunning Visualizations to help the growth of Business Profitability.
  • Proficient in writing functional specifications, translating business requirements to technical specifications, created/maintained/modified database design document with detailed description of logical entities and physical tables.
  • Proficient in writing functional specifications, translating business requirements to technical specifications, created/maintained/modified database design document with detailed description of logical entities and physical tables.
  • Extensive experience in Text Analytics, developing different Statistical Machine Learning solutions to various business problems and generating data visualizations using R, Python, and Tableau.
  • Experience in transforming business requirements into analytical models, designing algorithms, building models, developing data mining and reporting solutions that scales across massive volume of structured and unstructured data.
  • Expertise in transforming business requirements into analytical models, designing algorithms, building models, developing and reporting solutions that scale across a massive volume of structured and unstructured data.
  • Experience in designing stunning visualizations using Tableau software and publishing and presenting dashboards, Storyline on web and desktop platforms.
  • Hands on experience in implementing LDA, Naïve Bayes and skilled in Random Forests, Decision Trees, Linear and Logistic Regression, SVM, Clustering, neural networks good knowledge on Recommender Systems.
  • Strong knowledge of statistical methods (regression, time series, hypothesis testing, randomized experiment), machine learning, algorithms, data structures and data infrastructure.
  • Extensive hands - on experience and high proficiency with structures, semi-structured and unstructured data, using a broad range of data science programming languages and big data tools including Python, Spark, SQL, Scikit Learn, Hadoop MapReduce.
  • Proficient in Statistical Modeling and Machine Learning techniques (Linear, Logistics, Decision Trees, Random Forest, SVM, K-Nearest Neighbors, Bayesian, XGBoost) in Forecasting/ Predictive Analytics, Segmentation methodologies, Regression-based models, Hypothesis testing, Factor analysis.
  • Worked and extracted data from various database sources like Oracle, SQLServer, DB2, and Teradata.
  • Well experienced in Normalization &De-Normalization techniques for optimum performance in relational and dimensional database environments.
  • Extensive experience using data cleansing techniques, Excel pivot tables, formulas, and charts. Hand on working experience in machine learning and statistics to draw meaningful insights from data. I am good at communication and storytelling with data.
  • Highly motivated team player with excellent Interpersonal and Customer Relational Skills, Proven Communication, Organizational, Analytical, Presentation Skills, and Leadership Qualities.

TECHNICAL SKILLS

Programming Languages: Python, SQL, MATLAB, Spark, T-SQL, PL/SQL

Operating System: Windows 7/10, Unix, Linux

BI and Big Data Tools: Hadoop Map Reduce, Hive, Pig, Sqoop, Spark, Excel, PyCharm, Eclipse, Advanced Excel, Crystal Reports

Data Visualization: Tableau, ggplot2, Matplotlib, Seaborn, Plotly, SSRS

Database: Oracle, SQL Server, MySQL, PostgreSQL, MongoDB, NoSQL

ML / Deep Learning: Linear regression, SVR, KNN, Logistic regression, SVM, LDA, Random Forest, K-means, NLP, CNN, GAN, Autoencoder, Keras API

Python Libraries: Scikit, Pandas, NumPy, SciPy, Theano, Keras, Matplotlib, Seaborn, Plotly, Tensor Flow, Theano

Version Control: Git, GitHub, SVN

Reporting Tools: Tableau, SSRS

Environment: Jupyter, Anaconda, Spyder, Python Console, PyCharm

PROFESSIONAL EXPERIENCE

Confidential, Fairfax, VA

Data Analyst/ Data Scientist

Responsibilities:

  • CIPIO is the industry’s first platform to unite 360-degree customer behavioral insights with community influencers and their followers’ behavioral insights. The platform's first-of-its-kind precision influencer search tool pre-screens results so that marketers uncover the most real, relevant, and reachable community influencers for their brand. The project was to perform Data Profiling to learn about user behavior and merge data from multiple data sources.
  • Created dashboards, Data Mining, Data extraction, Data cleansing and Transformation, Prepared Reports & Insight from collected data using different Analytical functionalities.
  • Implemented ETL strategies for processing data while working with key users to derive report, data visualization, extraction maps/data mapping using Excel.
  • Performed Unit tests and Integration/acceptance tests for performance improvements and higher scalability of the application.
  • Implemented big data processing applications to collect, clean and normalization large volumes of open data usingHadoopecosystems such asPIG,HIVE, andHBase.
  • Worked on Clustering and classification of data usingmachine learning algorithms. UsedTensor Flowmachine learning to createsentimentally and time series analysis.
  • Implemented Classification using supervised algorithms Like Logistic Regression, Decision trees. KNN,Naïve Bayes.
  • Conducted analysis of assessing customer consuming behaviors and discover the value of customers withRMFanalysis; applied customer segmentation with clustering algorithms such asK-Means ClusteringandHierarchical Clustering.
  • Work on outlier’s identification withbox-plot, K-meansclustering using Pandas,NumPy.
  • Participate in features engineering such as feature intersection generating, feature normalize and Label encoding withScikit-learn pre-processing.
  • Use Python 3.0 (NumPy, SciPy, pandas, Scikit-learn, Seaborn,NLTK) andSpark 1.6 / 2.0 (PySpark, MLlib)to develop avariety of models and algorithms for analytic purposes.
  • Coordinate the execution ofA/Btests to measure the effectiveness of personalized recommendation system.
  • Perform data visualization with Tableau and generate dashboards to present the findings.
  • Recommend and evaluate marketing approaches based on quality analytics of customer consuming behavior.
  • Determine customer satisfaction and help enhance customer experience usingNLP.
  • Work on TextAnalytics,Naive Bayes, Sentiment analysis, creating word clouds and retrieving data from Twitter and other social networking platforms.
  • UseGitto apply version control. Tracked changes in files and coordinated work on the files among multiple team members.

Environment: MongoDB, Python, Spark (MLlib, PySpark), Tableau, Git, Unix, MLlib, Tensor Flow, Hadoop, HDFS, NLTK, MapReduce, OpenCV

Confidential, Boston, MA

Data Analyst/ Data Scientist

Responsibilities:

  • Confidential is a bioscience company innovating precision medicine through diagnostic biomarkers.
  • LSH focuses on the root of female biology, and we aim to offer the most effective, accessible, easy-to-use and affordable early-detection technology for disease in females of menstruating age.
  • Setup storage and data analysis tools in Amazon Web Services cloud computing infrastructure.
  • Used pandas, NumPy, Seaborn, SciPy, matplotlib, sci-kit-learn, NLTK in Python for developing various machine learning algorithms.
  • Worked on different data formats such as JSON, XML and performed machine learning algorithms in Python.
  • Participated in all phases of datamining; data collection, data cleaning, developing models, validation, visualization and performed Gap analysis.
  • Implemented Agile Methodology for building an internal application.
  • Programmed a utility in Python that used multiple packages (SciPy, NumPy, pandas)
  • Implemented Classification using supervised algorithms like Logistic Regression, Decision trees, KNN, Naïve Bayes.
  • Data transformation from various resources, data organization, features extraction from raw and stored.
  • Researched, evaluated, architected, and deployed new tools, frameworks, and patterns to build sustainable Big Data platforms for the clients
  • Identifying and executing process improvements, hands-on in various technologies such as Oracle, Informatica.

Environment: Python, AWS, JSON, SAS, NLTK, SQL Server, Oracle, MS Office, Erwin, Tableau, Hive, Hadoop, HDFS, PIG, MapReduce

Confidential, Boston, MA

Data Analyst/ Data Scientist

Responsibilities:

  • This project was for a retail client who had presence through online and offline retail stores across the country. The client wanted to understand the shopping behavior of its customers when showcasing new products or changing displays of the items it wants to sell fast based on certain criteria like visiting in specific section and previous shopping interests of the customers.
  • Conducted in-depth data analysis and predictive modeling to uncover hidden patterns and communicate the insights to the product, sales, and marketing teams.
  • Dealt with customer segmentation in determining the customer orientation based on price and category of the retails products, age group, ethnicity, visiting alone or with family and others using k -means clustering; Carried out sales forecasting using Regression Analysis with an accuracy of 85%.
  • Strategize cross functionally with marketing, product and finance on new product features and forecasting.
  • Identified levers to improve customer experience using supervised and unsupervised models on various attributes and levels within customer feedback data.
  • Application of Statistical and Analytical skills to derive actionable insights on the different Targeting population for different campaigns conducted by the Global Marketing Analytics team.
  • Building Models to satisfy the different Marketing Strategies and identifying the right audience so as to acquire new customers for the Consumer Products or increase the engagement/response of the existing customers to the Product.
  • Data extraction, data cleaning, exploratory data analysis, data transformations, data modelling and data visualizations using, SQL and Tableau
  • Wrote SQL queries using Common Table Expressions, Case statements, Set operators, Date formats and other DML statements
  • Performed data wrangling, manipulating functions along with customized user-defined functions
  • Developed models to assess trends, project needs, surface challenges, and estimate costs
  • Automated processes which reduced 50% of manual resources and doubled the efficiency
  • Used Tableau for creating dashboards, parameters to define upper and lower threshold, action and URL filters for navigation, table calculations and tableau inbuilt functions to provide effective solutions
  • Worked extensively on Advanced Analytics using LOD expressions, Scatter plots, Box and Whisker plots, Background images, Heat Maps, Pareto charts, Trend Lines and Log Axes, groups, hierarchies and sets, filters, to create detail level summary report and Dashboard using KPI's
  • Develop, maintain and document highly visualized comprehensive dashboards using Tableau Desktop and publish it to Tableau Server to meet business specific challenges
  • Developed dashboards using calculated fields with different logics for trends, parameters, calculations, groups, sets and hierarchies in Tableau

Environment: Python, Oracle, Tableau, MS Excel, Advanced Excel

Confidential, Sunnyvale, CA

Data Analyst/ Data Scientist

Responsibilities:

  • Confidential is a Technology Solutions company specializing in providing outsourced solutions. The project was for a retail client; Performed data cleansing and data manipulation for data preparation, performed ad-hoc reporting and exploratory data analysis to identify customer shopping patterns, analyze and interpret data for its eCommerce website.
  • Implemented end-to-end systems for Data Analytics, Data Automation and integrated with custom visualization tools using Hadoop, and MongoDB.
  • Gathering all the data that is required from multiple data sources and creating datasets that will be used in the analysis; collected marketing data to analyze market trend for strategic decision making
  • Performed Exploratory Data Analysis and Data Visualizations using Tableau.
  • Perform market analysis to efficiently achieve objectives of Sales optimization target, Pricing decision
  • Investigate and conduct studies on the forecasts, demand, and capital of products
  • Perform post campaign analysis and report valuable insights in EXCEL like possible reasons that result in the growth/decline of the campaign results.
  • Provide insights on consumer behavior and other key factors that could result in desired outcomes from the campaigns.
  • Analyze the market data and reports to provide the valuable insights that would help the marketing team achieve their goals.
  • Worked with Data Governance, Data quality, data lineage, Data architect to design various models and processes.
  • Developed, Implemented &Maintained the Conceptual, Logical & Physical Data Models using Erwin for forwarding/Reverse Engineered Databases.
  • Established Data architecture strategy, best practices, standards, and roadmaps.
  • Take up ad-hoc requests based on different departments and locations
  • Used Hive to store the data and perform datacleaning steps for huge datasets.
  • Created dash boards and visualization on regular basis using ggplot2 and Tableau.
  • Creating customized business reports and sharing insights to the management.
  • Interacted with the other departments to understand and identify data needs and requirements and work with other members of the IT organization to deliver data visualization and reporting solutions to address those needs.

Environment: Python, SQL Server, Oracle, MS Office, Tableau

Confidential, San Francisco, CA

Data Analyst/ Data Scientist

Responsibilities:

  • This project was focused on customer segmentation based on statistical modelling and predictive analystics; Responsibilities involved utilizing different data mining tools &statistical software to track and analyze data of custmers and help the management in decision making in driving business growth, develop a pricing model for various product and services bundled offering to optimize and predict the gross margin,
  • Worked with sales and Marketing team for Partner and collaborate with a cross-functional team to frame and answer important data questions prototyping and experimentation ML/DL algorithms and integrating into production system for different business needs
  • Optimized data collection procedures and generated reports on a weekly, monthly, and quarterly basis.
  • Used advanced Microsoft Excel to create pivot tables and pivot reporting, as well as use VLOOKUP function
  • Conducting the meetings with business users to gather data warehouse requirement
  • Gathered requirements and created use case diagrams as part of requirements analysis.
  • Performed data manipulation, data preparation, normalization, and predictive modelling. Improve efficiency and accuracy by evaluating model in Pytho.
  • Worked on Multiple datasets containing two billion values which are structured and unstructured data about web applications usage and online customer surveys
  • Performed Data cleaning process applied Backward - Forward filling methods on dataset for handling missing values.
  • Design, built and deployed a set of python modelling APIs for customer analytics, which integrate multiple machine learning techniques for various user behavior prediction and support multiple marketing segmentation programs.
  • Segmented the customers based on demographics using K-means Clustering.
  • Explored different regression and ensemble models in machine learning to perform forecasting.
  • Used classification techniques including Random Forest and Logistic Regression to quantify the likelihood of each user referring.
  • Performed Boosting method on predicted model for the improve efficiency of the model.
  • Designed and implemented end-to-end systems for Data Analytics and Automation, integrating custom, visualization tools using Tableau.

Environment: MS SQL Server, Python, Excel, Tableau, T-SQL, Access, XML, MS office, Oracle.

Confidential, El Dorado Hills, California

Data Analyst

Responsibilities:

  • The project was to develop Internet traffic scoring platform for ad networks, advertisers, and publishers (rule engine, site scoring, keyword scoring, lift measurement, linkage analysis).
  • Responsible for defining the key identifiers for each mapping/interface.
  • Implementation of Metadata Repository, Maintaining Data Quality, Data Cleanup procedures, Transformations, Data Standards, Data Governance program, Scripts, Stored Procedures, triggers and execution of test plans.
  • Used web crawling and text mining techniques to score referral domains, generate keyword taxonomies, and assess commercial value of bid keywords.
  • Developed new hybrid statistical and data mining technique known as hidden decision trees and hidden forests.
  • Reverse engineering of keyword pricing algorithms in the context of pay-per-click arbitrage.
  • Coordinated meetings with vendors to define requirements and system interaction agreement documentation between client and vendor system.
  • Automated bidding for advertiser campaigns based either on keyword or category (run-of-site) bidding.
  • Creation of multimillion bid keyword lists using extensive web crawling. Identification of metrics to measure the quality of each list (yield or coverage, volume, and keyword average financial value).
  • Enterprise Metadata Library with any changes or updates.
  • Document data quality and traceability documents for each source interface.
  • Establish standards of procedures.
  • Generate weekly and monthly asset inventory reports.

Environment: Python, SQL Server, Oracle, MS Office, Tableau

We'd love your feedback!