We provide IT Staff Augmentation Services!

Data Analyst Resume

4.00/5 (Submit Your Rating)

Boston, MA

SUMMARY:

  • 5.5+ years of experience in data analytics
  • Hands on experience in statistical techniques for Data Modelling, Data Architecture and Data Analysis
  • Sound noledge of Object - Oriented Programming concepts
  • Gleaning insights leveraging computational tools, and techniques (Statistical analysis)
  • Strong skills in Statistics Methodologies such as Hypothesis Testing, Principal Component Analysis (PCA), promotion TEMPeffectiveness and Mix modelling.
  • Skilled inPythonmodules like Pandas, NumPy, Matplotlib, Scikit - learn, PySide,SciPy, PyTables and generating complex graphical data
  • Outstanding preeminence in Data extraction, Data cleaning, Data Loading, Statistical Data Analysis, Exploratory Data Analysis, Data Wrangling, Predictive Modeling using R, Python and Data visualization.
  • Profound noledge in Machine Learning Algorithms like Linear, Non-linear and Logistic Regression, SVR, Reinforcement Learning, Natural Language Processing, Fuzzy Logic, Random forests, Ensemble Methods, Decision tree, Gradient-Boosting, K-NN, SVM, Naïve Bayes, Clustering (K-means), Deep Learning.
  • Proficient with Python 3.x including Numpy, Scikit-learn, Pandas, Matplotlib and Seaborn.
  • Excellent proficiency in model validation and optimization with Model selection, Parameter tuning and K-fold cross validation.
  • Intermediate Level R analytics expertise (Exploratory analysis using base graphs, statistical and Hypothesis testing).
  • Workedwith variousPythonIntegrated Development Environments like JupyterLab,PyCharm, PyScripter, Spyder, Visual Studio, PyDev and PyStudio
  • Basic understanding on Neural Networks, Deep Neural Networks and LSTM using TensorFlow and Keras.
  • Experience in using various version control systems like CVS, SVN and Git
  • Hands on experience in data analytics using R and Python
  • Expertise in supply chain management area including complex trials globally.
  • Hands on experience in documenting technical reports and project related documents for future s
  • Robust participation for functioning in fast-paced multi-tasking environment both independently and in teh collaborative team. Adequate with challenging projects and work in ambiguity to solve complex problems. A self-motivated exuberant learner.

TECHNICAL SKILLS:

Programming and Querying Languages: Python, R language, Bash scripting, SAS Base, Regular Expressions and SQL (MySQL&SQL Server).

Packages and tools: Pandas, NumPy, SciPy, Scikit-Learn, NLTK, ggplot2, dplyr, data.table,Keras, TensorFLow and PySpark

Machine learning: Linear Regression, Logistic Regression,Decision trees,Ensembles - Random Forest,Gradient Boosting, Xtreme Gradient Boosting(xGBM) and Support Vector Machines, Time series forecasting and Dimensionality Reduction

Reporting and Visualization: Matplotlib, Seaborn, ggplot2, iplots, PyTables and PySide

PROFESSIONAL EXPERIENCE:

Confidential, Boston, MA

Data Analyst

Responsibilities:

  • Collaborated with data engineers and operation team to implementETL process, wrote and optimizedSQL queriesto performdata extractionto fit teh analytical requirements.
  • Supervised data collection and reporting. Ensured relevant data is collected at designated stages, entered into appropriate database(s) and reported appropriately. Monitored assignments to assure distribution of workload and teh assessment of collection efforts.
  • Dealt with descriptive statistical analysis like customer profiling, content classifications and closed loop marketing data from databases.
  • Built predictive models including SupportVector Machine, Decision tree,Naive Bayes Classifier, Neural Networkto predict whether teh thyroid cancer cell is under potential danger of spreading by usingpython Scikit-learn.
  • Conventionally designed Fuzzy Logics and implemented statistical tests includingHypothesis testing, ANOVA, Chi-square testto verify models' significance by using R.
  • Participated in features engineering such as feature generating,PCA, feature normalizationandlabel encodingwithScikit-learn preprocessing. Data Imputation using variant methods inScikit-learn package in Python.
  • Work with business stakeholders to refine and respond to their ad hoc requests and improve their existing reporting and dashboards as necessary.

Environment: Python 3.x, R (ggplot2/ caret/ trees/ arules), Machine Learning (Logistic regression/ Random Forests/ KNN/ K-Means Clustering/ Gaussian Mixture Model / Hierarchical Clustering/ Ensemble methods/ Collaborative filtering), GitHub.

Confidential, NJ

Data Analyst

Responsibilities:

  • Communicated and coordinated with other departments to gather business requirements.
  • Gathering all teh data dat is required from multiple data sources and creating datasets dat will be used in analysis.
  • Responsible for end-to-end supply chain of Loreal products in performing teh data analysis activities.
  • Performed Exploratory Data Analysis and Data Visualizations usingR, andPython.
  • In Preprocessing phase, used Pandas andScikit-Learnto remove or impute missing values, detect outliers, scale features, and applied feature selection (filtering) to eliminate irrelevant features.
  • UsedPython(NumPy, Scipy, Pandas, Scikit-Learn, Seaborn) to develop variety of models and algorithms for analytic purposes.
  • Utilized spark, Matplotlib, Python, a broad variety of machine learning methods including classifications, regressions, dimensionally reduction etc.
  • UsedNLTKin Python for developing various machine learning algorithms.
  • Installed and usedCaffeDeep Learning Framework.
  • Modified selected machine learning models with real-time data inSpark (PySpark).
  • Worked with architect to improve cloud Hadoop architecture as needed for Research.
  • Worked on differentformats such asJSON, XMLand performed machine learning algorithms inPython.
  • Processed teh spend and goals data in Alteryx in such a way dat it is suitable for reporting
  • Participated in all phases of datamining; data collection, data cleaning, developing models, validation, visualization and performed Gap analysis.
  • UsedPandaslibrary for statistical Analysis.
  • Communicated teh results with operations team for taking best decisions.
  • Collected data needs and requirements by Interacting with teh other departments.

Environment: Python 3.2, MySQL, HTML, Python 2.7, XML, MySQL, MS SQL Server 2008/2012,Jupyter Notebook,RNN, ANN.

Confidential, IN

Data Analyst/Python

Responsibilities:

  • Investigated market sizing, competitive analysis and positioning for product feasibility.
  • Conducted research on development and designing of sample methodologies and analyzeddatafor pricing of client's products.
  • Responsible for all supply chain logistics activities supporting teh drug supply chain activities including raw materials
  • Worked on Business forecasting, segmentation analysis andDatamining.
  • DevelopedMachineLearningalgorithm to diagnose blood loss.
  • Generated graphs and reports using ggplot2 package inR-Studio for analytical models.
  • Developed and implementedRand Shiny application which showcases machine learning for business forecasting.
  • Developed predictive models using Decision Tree, Random Forest and Naïve Bayes.
  • Later usedAlteryxto blend thedata.
  • Performed analysis usingJMP.
  • Perform validation onmachinelearningoutput fromR.
  • Written connectors to extractdatafrom databases.

Environment: R,Python 2.x, Excel 2010, MachineLearning, Quick View, JMP, Segmentation analysis

Confidential

Senior Research Associate

Responsibilities:

  • Profiled trials from various registries belonging to USA, Europe, India, Japan, and Germany.
  • Collected and abstracted data from various trials including results.
  • Drafted clinical trial reports from various conference journals and press releases.
  • Updated daily Events News, Conferences, and arranged as client requirement.
  • Assembled research reports through secondary research, archival data and study audits.
  • Responded to Clients' queries, recorded weekly study reports and track data quality.

Environment: Excel 2010, CT.gov, EudraCT, Quick View, MySQL and SQL server

We'd love your feedback!