We provide IT Staff Augmentation Services!

Sr. Data Analyst Resume

2.00/5 (Submit Your Rating)

IL

SUMMARY

  • Over 7+ years of professional experience in Software Development, Database Development, Data Modelling, Data Analysis and System Engineering
  • Data Enthusiast with broad experience in IT and Data Technology oriented solutions with extensive knowledge of SDLC and Data Modelling.
  • Strong understanding of Spark Core, spark - SQL, PySpark, Spark Streaming and Machine Learning (SVM, Linear and Logistic Regression, KNN, Decision Tree, Random Forest, Gradient Boosting, Naïve Bayes and Cross Validation).
  • Strong Database Experience on SQL Server 2008 R2/2017 with T-SQL programming skills in creating Stored Procedures, Functions, Triggers and Views.
  • Experience in performing Exploratory Data Analysis (EDA), Dimensionality Reduction methods (PCA), missing value treatment and outlier treatment.
  • Adept at collecting, analyzing and interpreting large datasets to create visual insights and predictive models to meet the task requirements with creative problem solving and communication skills.
  • Expertise in using tools like Python, R, Tableau, big data platforms like Apache Hadoop, Apache Spark, Apache Kafka, Hive and writing and optimizing SQL queries.
  • Experience in NoSQL Column-Oriented Databases like HBase and MongoDB
  • Skills for debugging application code and problem solving for various production issues.
  • Hands-on expertise in predictive and prescriptive analytics - supervised and unsupervised machine learning models like regression and classification techniques including Linear Regression, Multiple Linear Regression, Logistic Regression, KNN, SVM, Random Forest, DBSCAN and Time-series data
  • Experience in Python to manipulate data for data loading and extraction and worked with python libraries like Matplotlib, Scikit, Numpy, Seaborn, TensorFlow, Keras and Pandas for data analytics and predictive modelling along with BI tool such as Tableau
  • Developed Natural Language Processing (NLP) pipeline: includes sentiment analysis, entities creation, and updating language model and also used NLTK in python.
  • Strong skills in Data Cleaning, Data Processing, Data Extraction & Visualization with large datasets of structured and unstructured data and Advanced Statistical Modeling, Correlation, Multivariate Analysis and ETL methodologies using SQL for validating and retrieving data from Hadoop
  • Used Tableau Desktop for the Data Analysis, Visualizations and for the explorations.
  • Experienced in creating and implementing visual KPI reporting tool for forward operations team leading to increased awareness of metrics.
  • Extensive experience in implementing the ETL processes which include developing the data model, designing the change capture and disaster recovery strategies.
  • Experienced and fully engaged in Software Development Life Cycle (SDLC) which includes gathering and analysing business requirements, functional/technical specifications, designing, developing, testing, deploying the applications and providing production support.
  • Expertise in performance tuning of Tableau reports by following best practices in source data like data blending, custom SQL etc.
  • Highly motivated team player with excellent interpersonal and customer relational skills, Proven communication, organizational, analytical, presentation skills, and leadership qualities.
  • Strong analytical, logical and problem-solving skills and ability to quickly adapt to new technologies by self-learning.

TECHNICAL SKILLS

Modelling Techniques: Decision Trees, Random Forests, Linear and Logistic Regression, Support Vector Machine, Naïve-Bayes, K-Nearest Neighbor, Neural Networks, Dimensionality Reduction, Clustering, NLP, Ensemble, Stack

Languages: Python, Java, Cypher, SQL, MATLAB, HTML, CSS, NoSQL, T-SQL, JavaScript, R

Machine Learning: Keras, scikit-learn, NLTK, TensorFlow

Hadoop Ecosystem: Hadoop, Spark, MapReduce, Pig, Flume, Hbase, Oozie, HDFS, Kafka, Hive, Cloudera

Data Analysis: Pandas, numpy, statsmodel, pyspark, sqlite3Data Visualization Tableau, Microsoft Excel, bokeh, plotly, matplotlib, seaborn

Others: Git, GitHub, GitLab, Jenkins, Docker, Linux

Databases: MySQL, MS SQL Server, PostgreSQL, MongoDB, Cassandra, Neo4J

PROFESSIONAL EXPERIENCE

Confidential, IL

Sr. Data Analyst

Responsibilities:

  • Extracted data from various sources for analyses, collaborated with subject matter experts to understand data artifacts.
  • Perform ETL (Extract, Transform and Load) data manipulation of large dataset. Identify and explore best suitable machine learning model by implementation and performance analysis of existing classification algorithms.
  • Performed exploratory data analysis and feature engineering on large datasets.
  • Created interactive visualizations using Tableau to present important key findings, parameters to business stakeholders
  • Developed various Python scripts to find vulnerabilities and data validation.
  • Building analytic models using a variety of techniques such as google analytics, logistic regression, risk scorecards and pattern recognition technologies
  • Experience in Python to manipulate data for data loading and extraction and worked with Python libraries like Matplotlib, Scikit-Learn, Numpy, Seaborn, TensorFlow, Keras and Pandas for data analysis.
  • Understanding in UNIX Shell scripts and writing SQL Scripts for development, automation of ETL process, error handling, and auditing purposes
  • Data Mapping, logical data modeling, created class diagrams and ER diagrams and used SQL queries to filter data within the Oracle database.
  • Created dashboards using Tableau Dashboard & prepared user stories to create compelling dashboards to deliver actionable insights.
  • Created extracts, published data sources to tableau server, refreshed extract in Tableau server from Tableau Desktop.
  • Involved extensively for building the dashboards such as creating Tableau extracts, refreshing extracts, Tableau layout designing, Tableau Work Sheet Actions, Tableau Functions, Tableau Connectors (Live and Extract), Dashboard color coding, formatting and report operations sorting, filtering, ranking, Top-N Analysis, hierarchies.
  • Worked on querying data and creating on-demand reports using Tableau Desktop and publish the same to Server.
  • Experience with Scheduling Tableau extracts, working with multiple data sources using data joining and data blending
  • Combined visualizations into interactive dashboards and publish them to the web.
  • Prototyped data visualizations using charts, drill-down, parameterized controls using Tableau to highlight the value of analytics in Executive decision support control.

Environment: R, Python, MATLAB, ETL, Sypder 3.6, Agile, Data Quality, R Studio, Tableau, Supervised & Unsupervised Learning, Java, NumPy, SciPy, h2o, Pandas, PL/SQL, RMSE, Matplotlib, Scikit-Learn.

Confidential, NY

Data Analyst

Responsibilities:

  • Connected Tableau server to publish dashboard to a central location for portal integration.
  • Connected Tableau server with share-point portal and setup auto refresh feature.
  • Created visualization for logistics calculation and department spend analysis.
  • Analyzed user and business requirements attended periodic meetings for changes in the application requirements and documents.
  • Understanding existing business model and customer's new requirements.
  • Created workbooks and dashboards using calculated metrics for different visualization requirements.
  • Data collection procedure enhancement to include information that is relevant for building analytic systems processing, cleansing, and verifying the integrity of data used for analysis
  • Optimized ETL processes by automating python scripts and reducing the overall execution time by 20%
  • Built analytical models using a variety of techniques such as google analytics, logistic regression, risk scorecards and pattern recognition technologies
  • Thorough grounding in all phases of data analysis, including definition and analysis of questions with respect to available data and resources, overview of data and assessment of data quality, selection of appropriate models and statistical tests and presentation of results.
  • Variable Identification, Missing value treatment, Outlier treatment, Variable transformation, Univariate and Bi-variate analysis.
  • Connected Hive tables with Tableau and performed data visualization for report.
  • Plot the trend and pattern Analysis and compare companies market capitalization from historical data.
  • Queried and retrieved data from SQL Server database to get the sample dataset.
  • In the pre-processing phase, used Pandas to clean all the missing data, datatype casting and merging or grouping tables for the EDA process.
  • Used PCA and feature engineering, feature normalization and label encoding Scikit-learn pre-processing techniques to reduce the high dimensional data (>150 features)
  • In data exploration stage used correlation analysis and graphical techniques in Matplotlib and Seaborn to get some insights about job application data.
  • Designed, developed and maintained daily and monthly summary, trending and benchmark reports in Tableau Desktop.

Environment: Python, Tableau Java, Cypher, SQL, MATLAB, HTML, CSS, JavaScript, PySpark, SciKit Learn, Keras., MS Access, MS Excel, MS Visio, UML diagrams, Mainframes, SQL Server.

Confidential

Data Analyst

Responsibilities:

  • Analyze, prepare and summarize financial data for various levels of management.
  • Collected, cleansed for modelling and analysis of structured and unstructured data used for major business initiatives.
  • Automated Driver-Partner enrollment with faster background checks to increase productivity.
  • Analyzing business data to develop Key Performance Indicators and create dashboards to visualize trends.
  • Collect data for analysis and build statistical model to optimize our Driver-Partners routes and schedule, which increased booking times by 35%.
  • Performed A/B testing during the development of the application and website to increase user retention.
  • Analyzed social media data to re-target customers using personalized ads and generate leads.
  • Worked with leadership teams to implement tracking and reporting of operations metrics across global programs
  • Performed K-means clustering, Multivariate analysis and Support Vector Machines.
  • Worked on Natural Language Processing with NLTK module for application development for automated customer response.
  • Utilized machine learning algorithms such as linear regression, multivariate regression, Naive Bayes, Random Forests, K-means, & KNN for data analysis.
  • Worked with large data sets, automate data extraction, built monitoring/reporting dashboards and high-value, automated Business Intelligence solutions (data warehousing and visualization)
  • Performed data entry, data auditing, creating data reports & monitoring all data for accuracy
  • Performed data discovery and build a stream that automatically retrieves data from multitude of sources (SQL databases, external data such as social network data, user reviews) to generate KPI's using Tableau.
  • Wrote ETL scripts in Python/SQL for extraction and validating the data.
  • Create data models in Python to store data from various sources.
  • Interpreting raw data using a variety of tools (Python, R, Excel Data Analysis Toolpark), algorithms, and statistical/econometric models (including regression techniques, decision trees, etc.) to capture the bigger picture of the business.
  • Created and presented dashboards to provide analytical insights into data to the client
  • Translated requirement changes, analyzing, providing data driven insights into their impact on existing database structure as well as existing user data.
  • Worked on SQL Server, creating Store Procedures, Functions, Triggers, Indexes and Views using T-SQL.
  • Involved in creating database objects like tables, views, procedures, triggers, functions using T-SQL to provide definition, structure and to maintain data efficiently.
  • Created action filters, parameters and calculated sets for preparing dashboards and worksheets in Tableau.
  • Effectively used data blending feature in tableau and defined best practices for Tableau report development.
  • Created Business requirement documents and plans for creating dashboards.
  • Deep experience with the design and development of Tableau visualization solutions.
  • Created personalized monthly reports for the drivers to maximize their profits efficiently.

Environment: MLbase, Pyspark, MapReduce, regression, logistic regression, random forest, neural networks, NLTK, XML, MLLib, Git & Json, Python, T-SQL, SQL, PL/SQL,, MS Access, MS Excel, XML, Microsoft Visio, UML, Unix

Confidential

ETL/BI Application Support

Responsibilities:

  • Write complex SQL queries for validating the data against different kinds of reports.
  • Created tables, indexes and constraints.
  • Worked with Excel Pivot tables.
  • Developed data access queries/programs using SQL Query to run in production on a regular basis and assists end users with development of complex Ad Hoc queries.
  • Created Error and Performance reports on Jobs, Stored procedures and Triggers.
  • Managed Database files, transaction log and estimated space requirements.
  • Configured and monitored database application.
  • Performing data management projects and fulfilling ad-hoc requests according to user specifications by utilizing data management software programs and tools like Excel and SQL.
  • Managed ongoing activities like importing and exporting, backup and recovery.
  • Involved in extensive data validation by writing several complex SQL queries and involved in back-end testing and worked with data quality issues.
  • Created and managed the databases in development environment.

Environment: SQL, PL/SQL, MS Office, MS SQL Server, SQL Developer, SSRS, MS Word and Excel

We'd love your feedback!