We provide IT Staff Augmentation Services!

Owner Resume

3.00/5 (Submit Your Rating)

Atlanta, GA

TECHNICAL SKILLS:

  • Data scientist with strong technical expertise, business and leadership experience, and communication skills to drive high - impact business outcomes through data-driven innovations and decisions.
  • Strong knowledge of statistical methods (regression, time series, hypothesis testing, randomized experiments), machine learning techniques, algorithms, data structures and data infrastructure.
  • Extensive hands-on experience and high proficiency with structured, semi-structured and unstructured data, using a broad range of data science programming languages and big data tools.
  • In-depth experience in R, Python, Spark, PySpark, SQL, MongoDB, Scikit Learn, Hadoop, Amazon AWS, Microsoft Azure, REST APIs, Unix, LINUX, GIT, R Shiny & ShinyDashboard.

PROFESSIONAL BACKGROUND:

Owner

Confidential, Atlanta, GA

Responsibilities:

  • Implemented an automated ETL process using PySpark, python and boto3 AWS API, processing billions of rows of mobile user geo-location and behavioral data from diverse sources.
  • Developed and deployed a clustering algorithm to discern user behavioral patterns from geo-location data, and ran it successfully over hundreds of millions of data points across hundreds of thousands of individuals on a daily basis. Leveraged a combination of PySpark parallelism and scikit-learn machine learning libraries to run large-scale clustering in tens of minutes.
  • Developed user behavior modeling in PySpark using geo-location, point-of-interest (POI) categorization (sourced from 3rd party data providers), time-of-day and other contextual information.
  • Wrote substantial PySpark implementations for data transformation, aggregation and summarization for post-ETL and post-clustering analysis and insights.

Confidential

Retail Supplier

Responsibilities:

  • Exclusively used R and associated packages to write/re-write many thousands of lines of code to implement two separate recommendation engines.
  • Used a highly modular and functional approach to R programming to deliver an easy-to-maintain and automated system.
  • Brought the entire code base into source control under Git, and instilled processes to keep code up-to-date and sharable.
  • Implemented a prototype product recommendation model in PySpark (Spark 2.1 and MLLib) using frequent pattern mining APIs for large data sets, which were intractable on single server in R.

Confidential

Health Provider

Responsibilities:

  • Implemented quality checking of FASTQ files to compute numerous industry-standard metrics in Spark (PySpark) and benchmarked the implementation on Microsoft Azure’s HDInsight and Amazon AWS’s EMR platforms. Ran the code on hundreds of FASTQ files spanning many tens of terabytes of data.
  • Experimented with cutting-edge spark-based parallel implementations of gene sequence alignment on the cloud and made recommendations on their usability / readiness.
  • Wrote a spark (PySpark) implementation that processed publicly available gene reference data - primarily from the 1000genome project, and compiled a re-usable repository of reference data to be used in multiple analytical studies. Delivered a proof-of-concept use of this reference data in analyzing genotypic and phenotypic correlations among subjects of a COPD (chronic obstructive pulmonary disease) study.

Confidential

Health Provider

Responsibilities:

  • Implemented and ran large-scale statistical models and tests on R on many tens of millions of data points across thousands of subjects, and tens of thousands of genes/variants with data sourced from Oracle databases using SQL and ROracle APIs (cohort data mart Confidential and omics data mart Confidential ).
  • Delivered valuable insights useful for both researchers and healthcare providers through R Shiny applications deployed in Microsoft Azure cloud.
  • Used a variety of robust statistical techniques to ascertain gene expression significance in different groups of subjects - with and without cancer, deceased and non-deceased and others. Applied association rules mining techniques to gene variant data to find the most commonly co-occurring variants in the subject population for variants within and across genes.
  • Developed numerous Shiny applications that integrated analytical insights from clinical and genomics data into patient dashboards, enabling care providers and other stakeholders with instant information.

Data Scientist

Confidential, Atlanta, GA

Responsibilities:

  • Transformed the company’s rudimentary and ad-hoc approach to data insights and reporting to a streamlined, automated, algorithmic, statistically robust and highly visual strategic function.
  • Worked extensively with countless data sources ranging in variety from structured SQL tables to unformatted text files, and in size from few rows to tens of millions of rows and tens of gigabytes. Drove unique insights into patient payment and online engagement behavior using creative data engineering, machine learning and experimentation techniques.
  • Implemented a first-generation regression-based predictive model that pre-identifies consumers that will fail to pay their bills. The model is designed to allow for ~30% reduction in processed bills with ~97% accuracy (3% false positive rate).
  • Developed a HIPAA-compliant record linking solution to enable billing and online patient wallet experience for patients that have multiple encounters at the same provider. Implemented a creative and a highly effective approach to link consumer records across providers to enable unique patient- and guarantor-level data analytics, inference and insights.
  • Deployed a suite of self-service, web application portals based on R Shiny and ShinyDashboard to easily access, visualize and consume data analysis outputs on an on-demand and real-time basis.

Data Scientist

Confidential, Florence, KY

Responsibilities:

  • Developed a statistical approach to assess post-launch impact of a new customer service telephony system on call handle time and associate performance. Leveraged R and SQL to make inferences using observed data.
  • Designed and developed a regression model to predict call handle times per call and customer issue type using a new creative approach that successfully combined previously disjoint data sets. Sourced data from raw server logs, and existing databases, and used data engineering and cleaning techniques for subsequent use in the regression model.
  • Provided insights into the effectiveness of a recruitment aptitude test in hiring associates using pre- and post-treatment associate performance metrics. Relied on statistical tests to assess differences between the two populations across a wide range of metrics.

We'd love your feedback!