Principal Scientist Resume
5.00/5 (Submit Your Rating)
SUMMARY:
- Persistence in decoding life using numbers and symbols
- PhD in bioinformatics with over 10 years work experience in the bio - pharmaceutical industry, great software development and data analysis skills, is looking for a bioinformatics position.
HIGHLIGHTS OF EXPERTISE:
- Bioinformatics
- Asthma/Cancer Research
- Personalized Medicine
- Geno-Pheno relationship
- Software Engineering
- Certified Statistics Modeler
- Genetic Epidemiology
- NGS, DNAseq, RNAseq
- Bio-Pharmaceutical Industry
- Data Mining, Analysis
- Encode, HapMap, dbSNP, 1000G, HGMD, OMIM
COMPUTER LANGUAGES & SKILLS:
- JAVA Perl Python C/C++ HTML /HTML5 VB Pascal Assembly JavaScript XML
- CVS/SVN EJB J2EE/Struts Jsp/servlet UML
- JDBC/ODBC SQL PL/SQL CGI R Oracle Informix mySQL PostgreSQL Access
- UNIX Linux MS NT/WIN MAC Cloud Computing/AWS Android/iPhone/PhoneGap
- Fastq-dump Fastqc BWA Samblaster Samtools GATK Freebayes Hisat Stringtie
PROFESSIONAL EXPERIENCE:
Principal Scientist
Confidential
Responsibilities:
- Created a wide selection of genetic testing products, including personalized medicine, disease risk prediction personalized nutrition, and inborn talents, etc.
- Built a Chinese genome-wide phenotype prediction system with a proprietary geno-pheno database, a prediction algorithm, and computer programs (Perl, Java, R, HTML5, MS Excel)
- Integration of geno-pheno data sources from HapMap, SNAP, NHGRI, etc. (Perl)
- Internal QA of data (Perl)
- Database construction (Perl, MS Excel)
- Selection and implementation of geno-pheno assessment algorithm (Perl, Java, R)
- Data preprocessing (Perl)
- Design and implementation of a genetic test report system (Perl, HTML5)
- Supervised publication on ‘Validation of warfarin pharmacogenetic algorithms in 586 Han Chinese patients’
- Data collection, QA, analysis (Linear regression, R)
- Guidance on paper writing & partial paper writing
Postdoc Research Fellow
Confidential, Boston, MA
Responsibilities:
- Genome wide association studies ( Confidential )
- Pancreatic cancer survival
- Goal to identify genetic and environmental factors associated with pancreatic cancer survival, and to build prognosis models to predict survival of pancreatic cancer patients.
- Materials: pancreatic cancer patients from 12 cohorts with ~1200 subjects
- Results: partial results identified some risk factors with strong confidence.
- Publication: see publication.
- Contribution: collaboration with Confidential Medical School and National Cancer Institute. Prepared the Confidential part of the data.
- Mole count and melanoma risk
- Goal: to identify novel genetic variants associated with mole count.
- Methods: linear regression adjusted for age, gender, and Eigen Vectors. Betas from each study of the discovery set were combined by a meta-analysis with weights proportional to the inverse variance of the beta in each study.
- Goal: to identify genetic variants associated with, and to build a model to predict handedness.
- Methods: logistic regression adjusted for Eigen Vectors. Betas from each study of the discovery set were combined by a meta-analysis with weights proportional to the inverse variance of the beta in each study. Heritability estimation of handedness is computed using EMMAX.
- Contribution: basically did all the computational work and produced all the results. Involved in the experimental design.
- Perl for data processing and pipelining, SAS for data retrieval, R for statistics and plotting, shell scripts for pipelining.
- Disease risk prediction using environmental and genetic risk factors
- ABO blood type and pancreatic cancer risk
- Analyze the relationship between ABO blood type and pancreatic cancer risk.
- Results: An increased risk was observed in participants with A1 but not A2 alleles.
- ABO blood type and breast cancer risk and survival
- Analyze the relationship between ABO blood type and breast cancer risk and survival. used SNPs to infer ABO blood types (A1, A2, B, O1, O2) and secretor (a protein involved in the secretion of the ABO antigens) status.
- Perl for data processing and pipelining, R for statistical analyses.
Pre-doctoral Research Fellow
Confidential, Boston, MA
Responsibilities:
- Facilitated investigative research on bacteriology, chronic disease epidemiology, and virology as part of multidisciplinary research division of Brigham and Women’s Hospital, and Confidential Medical School
- Prepared Ph.D. thesis on ‘Computational Approach For The Identification Of Functional Regulatory Genetic Polymorphism’, using gene regulatory elements, DNA sequence features, and genetic markers
- Conducted leading-edge study on development of computational model to predict asthma exacerbation with thousands of SNPs, leveraging on clinical attributes, and 550k genetic markers, and Random Forests
- Published ‘Study to Predict Severe Asthma Exacerbations in Children using Random Forests Classifiers’
Sr. Software Engineer
Confidential, Berkeley, CA
Responsibilities:
- As the principle engineer in the middle tier group, designed, implemented, and successfully delivered Confidential 2.0, a microarray gene expression data management software product providing support for exploration and analysis of gene expression data for tens of thousands of biological samples, together with associated clinical data, gene annotation, and metabolic pathway information included in Confidential and Confidential . Confidential is Confidential ’s flagship product and is currently used by the top pharmaceutical companies.
- As a team member, helped requirements collection, functional specification, design documentation, and test case specification
- As the principle engineer in the middle tier group, involved in the design and implementation of GUI (MVC, OODA; Swing; Query tools and analysis tools)
- Completed the design and implementation of the middle tier that interfaces with the Oracle 9i, Corba and HTTP servers and the GUIs (MVC, OOAD; Java, C/C++, Corba, HTTP, XML; JUnit)
- Completed the design and implementation of workflow diagram engine that can significantly speed up general computing by caching (OOAD; Java, XML serialization and deserialization; JUnit)
- As a senior software engineer, helped the XML file migration from old versions to the new version (Perl, Java, XML)
- Spearheaded the use of design patterns (MVC, OOAD; TogetherJ) and formal software development process (Rational Unified Process or RUP)
- Strove for a higher software testing standard, pioneered and constructed a complete set of unit testing (JUnit) and functional testing (WinRunner).
Computational Scientist
Confidential, Seattle, WA
Responsibilities:
- Expression Gene Array Database ( Confidential ) System - a Web based application system to register gene array data for drug discovery R & D
- Technical lead
- Adopted the Rational Unified Process (UML, RequisitePro, Rose, ClearQuest, Robot) and technically managed the full software development lifecycle
- Tomcat/JSP/servlet, JavaScript, LiveConnect, J2EE/EJB, Struts; Oracle 8.1.5, PL/SQL; Solaris; CVS
- Chemical Registration ( Confidential ) System - a Web based application system to register and display chemical structures and related data for drug discovery R & D
- Server Side, Client Side JavaScript/LiveConnect for the Web front-end
- MDL PL for some back-end work
- Perl (mostly OO Perl) for data processing and batch registration
- Oracle 8.1.5, PL/SQL; DBI/DBD
- Rational Suite certificate and Access programming certificate
- Unified Drug Discovery Suite to integrate and manage the chemical data, high through put Screening data, and microarray data (Visual Basic, ODBC. Project cancelled after months of preparation)
- LIMS requirements collection, functional specification, and evaluation
- Designed and implemented an innovative incremental XML DOM parser
- EST clustering analysis to examine the ancient Chinese “5-element” theory
- SNPs analysis to study the relationship of mutation rate and DNA functional regions
- DNA sequence analysis to study the distribution of entropy over fixed length short sequences
