We provide IT Staff Augmentation Services!

Data Analyst Resume

4.00/5 (Submit Your Rating)

Nyc, NY

SUMMARY

  • Result - oriented data scientist wif 5+ years of IT experience along wif outstanding ability of developing solutions along wif communication and collaboration across teams
  • Experience in working in a fast-paced team to identify business problems, and determine teh data source required to fulfill requirements from clients’ and different function teams
  • Hands o experience handling different types of data and data cleaning, data manipulation and data transformation
  • Expertise in multiple programming languages, and data analytics & visualization tools including R, Python, SQL
  • Experienced in designing and maintaining relational databases using MySQL, MSSQL Server and PostgreSQL
  • Experience of SQL queries on data manipulation using Window functions and sub-queries
  • Experienced in data warehouse and ETL technologies
  • Solid at applying Machine Learning, Deep Learning models and Hyperprameter Tuning
  • Adapt knowledge of big data tools like Hadoop (HDFS, Hive, MapReduce) and Spark (SparkSQL, Spark MLlib)
  • Working knowledge of object and bucket operations on Amazon Web Services S3 and instance operations on Amazon Web Services EC2
  • Strong analytical and creative problem-solving skills as well as teh ability to make effective decisions
  • Experience in database schemas designed, data transformation from raw data to designed database schemas, and data cleansing procedures to extract data
  • Experienced in handling and manipulating NoSQL databases using MongoDB
  • Ability to communicate and maintain good working relations wif technical team members, and technical and non-technical partners and stakeholders
  • Strong organizational skills along wif teh ability to handle multiple tasks simultaneously
  • Hands on experience on project management tools such as Jira
  • Experience wif Web Scraping, Data Extraction, Cleansing, Visualization, Microsoft Excel PivotTable

TECHNICAL SKILLS

Statistical Modeling: Regression, Classification, Clustering, Predictive Modeling

Machine Learning: Advanced coursework in mathematics and statistics, scikit-learn, TensorFlow, PyTorch

Programming: proficient in Python (wif pandas, NumPy, matplotlib), R, and Microsoft Excel; familiar wif JAVA, HTML, SQL

Databases: MySQL, Oracle SQL, SQL Server, NoSQL, SAS, MongoDB

Data Visualization& BI: Tableau, PowerBI, Microsoft Excel

Platforms: prior experience wif Hadoop, Azure, Pig, Hive, Spark; Airflow, AWS s3, sqs, ec2, Kubernetes, Docker, Databricks, Google Cloud

Operating Systems: MacOS, Windows, Unix/Linux

PROFESSIONAL EXPERIENCE

Data Analyst

Confidential, NYC, NY

Responsibilities:

  • Worked closely wif engineers to deploy models in production and systematically track model performance, including creating and monitoring teh dashboards
  • Performed data exploration and analysis to determine coverage of new data, to evaluate correlation to key business metrics, to determine teh importance of various variables and provide recommendations on whether teh data source provides incremental value to Sirius
  • Assisted in engagements wif key business stakeholders in discussions on business strategies and opportunities
  • Prepared and presented executive summaries of findings and detailed working team communications, wif responsive iterations on such presentations based on individual data visualization preferences
  • Preprocessed teh review texts by Spark NLP Pipeline of Tokenizer, Lemmatizer and StopWordsCleaner
  • Analyzed teh keywords for product reviews and personalized product recommender based on PySpark ALS Algorithm Developed a matrix factorization-based
  • Transformed and automated database wif Kubernetes
  • Worked on Hadoop, Hive, Pig and Spark to manage data processing and storage big data applications running in clustered system
  • Applied Natural Language Processing models/tools to automatically analyze teh underlying structures of multiple prouct review datasets, in order to better understand teh customer feedback/concerms
  • Operated ETL (extract, transform and load) on NoSQL unstructured data
  • Conducted data visualization using Tableau in collaboration wif marketing team and product team

Environment: Python, Spark, Kubernetes, Hadoop, Hive, Pig, Tableau

Data analyst

Confidential, NYC, NY

Responsibilities:

  • Explored teh data of 9 million observations and performed data cleaning, filtering and merging using SAS and SQL, created potential customer lists under clients’ requirement
  • Recalibrated Decision Tree model wif updated data in Python, improved 5% of accuracy comparing to teh previous model
  • Worked closely wif developers and database administrators to transform data models from logical to physical
  • Generated dashboards wif story-telling quarterly business performance interactive dashboards for quarterly summary meeting wif Tableau
  • Demonstrated results using charts in Microsoft Excel
  • Predicted teh business outcomes such as order volumes by splitting data into training and testing datasets for regression model
  • Identified distinct groups of customers which informed content and marketing strategy

Environment: Annaconda, Jupyter Notebook, Python, SQL, SAS, Tableau, scikit-learn

Data Analytics

Confidential

Responsibilities:

  • Performed data cleansing and standardization using text mining techniques on data collected from Spectrum and VMware to mitigate discrepancies between two data sources during data merging
  • Identified high-risk servers by categorizing servers into groups and targeting those only exist in one of teh data sources
  • Collaborated wif a cross-functional team from diverse demographic backgrounds and won Innovation Award based on creativity and initiative

Environment: Python, VMware, Microsoft Excel

Data analyst

Confidential, Washington DC

Responsibilities:

  • Experimented wif various Machine Learning techniques to predict stock price based on teh influence of social media by investigating public sentiment, using Twitter API and text mining techniques, based on over 60,000 rows of data
  • Developed performance hypothesis tests to discover teh influential factors and conducted clustering analysis wif 3 different methods (Ward, K-means and DBSCAN) and evaluated each method, using Silhouette coefficient as a metric
  • Constructed prediction models including Support Vector Machine (SVN), Naïve Bayes (NB) and Random Forest (RF)
  • Compared teh accuracy and selected SVM as teh final approach to improve teh accuracy from 53% to 67%

Environment: Python, NLP, Web Scraping, Machine Learning, Classification

Data Analyst

Confidential, Charlotte, NC

Responsibilities:

  • Developed leader-follower algorithm to find teh connections between nodes wif R
  • Worked wif stakeholders and technology teams to understand system requirements
  • Partnered wif business and technology teams to document, identify and implement solutions to system, risks, defects and data deficiencies.

We'd love your feedback!