Research Classification
Research Interests
Relevant Thesis-Based Degree Programs
Affiliations to Research Centres, Institutes & Clusters
Research Options
Research Methodology
Recruitment
Our research interests include:
- Integration of metabolomics with other ‘omics’ (epigenomics, genomics, transcriptomics, and proteomics) data for the systems-level interrogation of biological problems
- Application of state-of-art metabolomics technologies in various biological challenges, such as mechanistic understanding of cancer metabolism and disease biomarker discovery
- Synergetic development of analytical and bioinformatic techniques to enhance metabolomic coverage and improve the confidence of metabolite identification
Complete these steps before you reach out to a faculty member!
Check requirements
- Familiarize yourself with program requirements. You want to learn as much as possible from the information available to you before you reach out to a faculty member. Be sure to visit the graduate degree program listing and program-specific websites.
- Check whether the program requires you to seek commitment from a supervisor prior to submitting an application. For some programs this is an essential step while others match successful applicants with faculty members within the first year of study. This is either indicated in the program profile under "Admission Information & Requirements" - "Prepare Application" - "Supervision" or on the program website.
Focus your search
- Identify specific faculty members who are conducting research in your specific area of interest.
- Establish that your research interests align with the faculty member’s research interests.
- Read up on the faculty members in the program and the research being conducted in the department.
- Familiarize yourself with their work, read their recent publications and past theses/dissertations that they supervised. Be certain that their research is indeed what you are hoping to study.
Make a good impression
- Compose an error-free and grammatically correct email addressed to your specifically targeted faculty member, and remember to use their correct titles.
- Do not send non-specific, mass emails to everyone in the department hoping for a match.
- Address the faculty members by name. Your contact should be genuine rather than generic.
- Include a brief outline of your academic background, why you are interested in working with the faculty member, and what experience you could bring to the department. The supervision enquiry form guides you with targeted questions. Ensure to craft compelling answers to these questions.
- Highlight your achievements and why you are a top student. Faculty members receive dozens of requests from prospective students and you may have less than 30 seconds to pique someone’s interest.
- Demonstrate that you are familiar with their research:
- Convey the specific ways you are a good fit for the program.
- Convey the specific ways the program/lab/faculty member is a good fit for the research you are interested in/already conducting.
- Be enthusiastic, but don’t overdo it.
Attend an information session
G+PS regularly provides virtual sessions that focus on admission requirements and procedures and tips how to improve your application.
ADVICE AND INSIGHTS FROM UBC FACULTY ON REACHING OUT TO SUPERVISORS
These videos contain some general advice from faculty across UBC on finding and reaching out to a potential thesis supervisor.
Graduate Student Supervision
Doctoral Student Supervision
Dissertations completed in 2010 or later are listed below. Please note that there is a 6-12 month delay to add the latest dissertations.
Chemical pattern recognition in metabolomics and exposomics by mass spectrometry and machine learning (2025)
Metabolomics and exposomics has gained increasing attention. They focus on profiling small molecules derived from biological processes and environmental sources. Recent studies reveal that genetic factors alone cannot explain the majority of chronic diseases, while environmental factors display close association with disease development and progression. Identifying novel and emerging pollutants is essential for uncovering toxicity drivers, regulating their levels, and minimizing exposure, ultimately contributing to improved human health. Liquid chromatography coupled with high-resolution mass spectrometry (LC-HRMS) is a prominent tool for screening the chemical composition and discovering unknown compounds in diverse samples. Among exposure chemicals, disinfection by-products (DBPs) generated during water disinfection treatment represent an important class. Amine-containing compounds are particularly reactive precursors in water sources, leading to the formation of odorous and toxic nitrogenous DBPs. Additionally, halogenated species—such as chlorine-, bromine-, and iodine-containing compounds—are persistent, bioaccumulative, and toxic, posing significant health risks. Despite more than 50 years of DBP research, over 50% remain unidentified. The identification of these compounds is challenging due to analytical variations, chemical diversity, trace-level concentrations, and a lack of available spectral databases. Furthermore, the high sensitivity and throughput of LC-HRMS-based omics generate large and complex datasets, presenting significant challenges in its data processing and interpretation.This Ph.D. thesis focuses on applying machine learning and computational techniques to enhance the data processing of LC-HRMS analyses and the identification of exposure chemicals. Chapter 1 provides an overview of the technologies and introduces research topics addressed in this thesis. Chapter 2 discusses the issue of contaminated tandem spectra and developed a machine learning platform that improves tandem spectral quality for structural annotation. Chapter 3 developed HDPairFinder to selectively identify amine-containing compounds in isotope labelled MS data. Chapters 4 and 5 address the limitation of current approaches for the identification of chlorine- and bromine-containing compounds. Utilizing machine learning and cloud computation, ChloroDBPFinder and HalogenFinder were constructed for compound discovery and classification. Chapter 6 aims for comprehensive screening of iodine-containing compounds in both positive and negative ionization modes. A standalone program, IodoFinder, was developed to utilize characteristic fragmentation patterns to achieve the discovery of iodine-containing compounds.
View record
Development of analytical methods and bioinformatic programs for data acquisition and data processing in LC-MS-based untargeted metabolomics (2023)
Untargeted metabolomics studies the complete set of small metabolic molecules in a given biological system. Liquid chromatography coupled with high-resolution mass spectrometry (LC-HRMS) is currently the most prominent analytical platform for untargeted metabolomics owing to its high sensitivity, specificity, and metabolic coverage. However, the current LC-MS-based untargeted metabolomics workflow has limited performance in detecting and quantifying trace-level metabolites of bad chromatographic peak shapes. It is also hard to differentiate signals of real metabolites from noise and background. During my Ph.D., I have developed a suite of analytical and bioinformatic tools to address the critical challenges in metabolomics data acquisition and data processing. In this thesis, Chapters 2 to 4 focus on the development of data acquisition methods. Specifically, Chapters 2 to 3 describe the detailed comparison of the existing data acquisition modes in two different aspects. Chapter 4 describes the development of a novel data acquisition strategy, DaDIA, to increase metabolomic coverage. On the other hand, Chapters 5 to 9 describe the development of data processing methods. In Chapter 5, a data processing parameter optimization tool, Paramounter is introduced to rapidly and accurately determine the best peak-picking parameters for five commonly used data processing programs. In Chapter 6, the five most commonly used metabolomics data processing programs were compared to mechanistically explain the difference regarding the performance in metabolic feature extraction. In Chapter 7, a novel data processing program, JPA, was developed to efficiently extract the metabolic features by combining multiple peak-picking algorithms. In Chapter 8, a deep learning-based software, EVA, was created to automatically remove the false positive features generated from the background noise. In Chapter 9, I developed a bioinformatic workflow, ISFrag, to automatically recognize and remove the false positive metabolic features originating from in-source fragmentation.
View record
Development of analytical workflows and bioinformatic programs for mass spectrometry-based metabolomics (2023)
Quantitative determination of metabolite concentrations in biological samples is fundamental to biological and clinical research. Metabolomics analyzes the entire set of metabolites in a given biological system. It is an emerging technology in the post-genomic era to interrogate cellular biochemistry, perform diagnostic testing, stratify patient populations, and characterize biochemical mechanisms of disease. Recent successes in metabolomics demonstrate the central role of mass spectrometry (MS) in small molecule quantification, owing to its high sensitivity, high throughput, and broad metabolic coverage. Even though diverse MS instruments have been developed for metabolite quantification, it is still challenging to quantify the entire metabolome accurately and precisely. Besides MS hardware advances, quantitative metabolomics also requires extensive efforts in other analytical and bioinformatic methodology development. For a given MS platform, analytical method development focuses on laboratory practice, including sample handling, metabolome extraction, and data acquisition. In comparison, bioinformatic method development emphasizes computational data processing, such as data calibration, data curation, and statistical analysis. The subsequent chapters detail the development of analytical and bioinformatic solutions for quantitative metabolomics from improving metabolic coverage, analytical accuracy, analytical precision, and statistical analysis. Lastly, this thesis describes a metabolomics study of mouse brain regional differences in metabolism between males and females. Collectively my studies of quantitative metabolomics improve quantitative performance, deepen our knowledge of the MS-based quantification process, and facilitate the generation of confident biological conclusions.
View record
Towards accurate compound annotation in mass spectrometry-based global metabolomics (2023)
Metabolomics is an emerging omics study that aims to characterize the entire metabolome in a biological system. Mass spectrometry (MS) is a preferred analytical technique for metabolomics research owing to its high sensitivity and highly specific structural information content. However, it remains a longstanding challenge to accurately translate MS signals into chemical language, thus hindering the downstream biological interpretation.This dissertation presents computational strategies contributing to tandem mass (MS/MS) spectral interpretations with the aid of machine learning and statistical approaches. Chapter 1 provides a holistic introduction to MS-based metabolomics and the developed bioinformatic tools for uncovering the unidentified metabolic features in untargeted metabolomics. Chapter 2 describes a novel MS/MS spectral comparison algorithm, Core Structure-based Search (CSS), which searches for structural analogs of unknown MS/MS spectra within the existing MS/MS reference libraries. CSS shows improved correlations with structural similarity in large-scale benchmarking. In Chapter 3, a deep learning-based tool is developed for automated extraction of steroid-like metabolic features from the untargeted metabolomics data by classifying MS/MS fragmentation patterns. This biology-driven metabolomics pipeline enables metabolite characterization and discovery on the compound class level. Chapter 4 depicts the purification of chimeric MS/MS spectra using a random forest model. Purified MS/MS spectra are demonstrated to yield better spectral matching results against MS/MS reference libraries. Chapter 5 describes the systematic analysis of radical fragment ions in MS/MS through MS/MS database mining. Larger than expected percentages of radical ions are present in collision- induced dissociation-based MS/MS; relationships between radical ion percentages and compound classes, chemical substructures and collision energies are also investigated. Chapter 6 discusses a standalone platform, BUDDY, for molecular formula discovery via bottom-up MS/MS interrogation and experiment-specific global peak annotation. BUDDY further integrates machine-learned ranking and significance control, showing improved formula annotation accuracy and lower computational cost than other benchmarking tools. Applying BUDDY on repository- scale recurrent unidentified MS/MS spectra, we discovered >5,000 chemical database-unarchived molecular formulae with high confidence. Overall, this dissertation demonstrates computational contributions to enriching structural insights into MS-based untargeted metabolomics data, thus paving the way for understanding biological mechanisms behind various health disorders and diseases from the perspective of small molecules.
View record
Master's Student Supervision
Theses completed in 2010 or later are listed below. Please note that there is a 6-12 month delay to add the latest theses.
Development of global metabolomics and its application in molecular understanding of extracellular vesicles in parasite infection (2026)
This thesis delves into metabolomics, the study of small molecules—known as metabolites—present within a biological system. These molecules span both polar classes, such as amino acids, sugars, and nucleotides, and non-polar classes, such as fatty acids and other lipid species, which are often collectively referred to as lipids. The study of this latter group is often treated as a distinct domain called lipidomics, which focuses specifically on the comprehensive analysis of lipids. Accordingly, the term “metabolomics” may be used either as a broad umbrella that includes both polar and non-polar molecules, or a narrower term that refers to the study of polar molecules only. This work adopts the latter definition and focuses on integrating metabolomics with lipidomics to enable a more comprehensive investigation of the biological systems of interest. The first project focuses specifically on lipidomics, examining the lipid composition of extracellular vesicles secreted by wild-type and drug-resistant Leishmania parasites. By annotating and comparing lipid species present in these vesicles, this study provides detailed insight into lipid remodeling associated with drug resistance. Extracellular vesicles play a critical role in parasite–host interactions, and alterations in their lipid composition may influence membrane properties, signaling processes, and resistance mechanisms. This lipid-centric approach enables high-resolution chemical profiling of resistance-associated changes and contributes to the identification of lipid signatures that may serve as biomarkers or therapeutic targets in leishmaniasis.The second project expands beyond a single molecular class and addresses a key methodological challenge relevant to both metabolomics and lipidomics: the comparison of single-phase and dual-phase extraction strategies. While dual-phase extraction enables the simultaneous recovery of polar metabolites and lipids from a single sample, it may result in reduced analyte concentrations compared to single-phase methods. This study systematically investigates three mechanistically distinct factors hypothesized to drive these differences—pipetting effect, partitioning effect, and matrix effect. By isolating and evaluating each factor, this work provides a framework for objectively assessing extraction performance and informs methodological decision-making in integrative small-molecule analyses.Overall, this thesis advances the field by combining a lipidomics-focused biological application with a broader methodological evaluation relevant to integrated metabolomics–lipidomics workflows.
View record
Development of computational solutions to process and interpret mass spectrometry-based metabolomics data (2024)
Metabolomics, the study of small molecules within biological systems, offers valuable insights into biochemical processes. This thesis addresses two challenges in mass spectrometry-based metabolomics: processing complex breathomics data and enhancing the structural annotation of metabolites using deep learning.In the first study, I introduce BreathXplorer, an open-source Python package designed to process real-time exhaled breath data from secondary electrospray ionization high-resolution mass spectrometry (SESI-HRMS). BreathXplorer tackles the challenge of non-Gaussian metabolic signal shapes by employing topological algorithms or Gaussian mixture models (GMM) to identify exhalation intervals and density-based spatial clustering (DBSCAN) to cluster m/z values. It accurately determines the start and end points of exhalation, ensuring precise quantitative measurements. In a proof-of-concept study on exercise breathomics, BreathXplorer identified exercise-responsive metabolites, showing its potential in real-time metabolomics research.In the second study, I explore deep learning as a solution for compound annotation in tandem mass spectrometry (MS/MS). This offers a predictive strategy where traditional library search tools are limited due to small spectral libraries. Generative models have become a fundamental approach for various tasks; however, their application to MS/MS structural annotation is less developed and often underperforms compared to traditional fingerprint-based counterparts. In this work, I investigate the potential of generative models by building and comparing transformer-based fingerprint models and generative models. This comparison helps to understand their strengths and limitations in annotating chemical structures. Training and testing on 616,594 unique structures, I identified three key limitations of direct structure generation with generative models: (1) error accumulation, (2) generation of invalid compounds, and (3) generation of unrecorded compounds. I then propose a solution using the generative model as a ranker. My results demonstrate that generative-based ranking outperforms fingerprint-based systems. Furthermore, I observed a positive correlation between high generation scores and the correctness of predicted structures. My analysis shows that generative models can leverage richer structural information during training, leading to improved accuracy in end-to-end chemical structure identification.
View record
Development of bioinformatic solution to enhance metabolomics data quality and its application in plant research (2023)
This thesis delves into the development and application of metabolomics, a discipline focused on the comprehensive study of metabolites within biological systems. The research is segmented into two interconnected parts: methodology development for metabolomics data processing and its subsequent application in plant stress physiology. The first project tackles the challenge of computational variation in untargeted metabolomics, which arises due to the incapability of data processing for complex LC-MS data. An in-depth exploration led to the identification of sources and causes of computational variation, followed by the development of novel methodologies to mitigate these challenges. These methodologies, including data processing parameter optimization and a machine learning program, successfully reduced computational variation, thereby enhancing the quantitative precision of untargeted metabolomics. The second segment of the thesis applies these methodologies to study the salinity stress response in Alfalfa (Medicago sativa L.). Comprehensive analysis of the plant's metabolic alterations, coupled with transcriptomics data, revealed significant pathways and mechanisms of salinity response. The integration of multi-omics data provided a deeper understanding of the complex interplay between genes and metabolites. The research advances the field of metabolomics, providing improved data processing methodology and valuable insights into plant stress physiology. Future work may expand these findings towards personalized medicine, disease diagnosis, and precision treatment.
View record
Development of bioinformatics solutions to enable hair-based exposome research (2023)
Metabolomics and exposomics are rapidly expanding fields that aim to understand the intricate relationship between environmental exposures and human health. Hair is an underexplored matrix for studying metabolomics and exposomics, serving as a record of chemicals deposited on its surface. However, the lack of a comprehensive database and compound annotation pipeline has hindered the use of hair in untargeted studies. To address these challenges, a comprehensive database of hair metabolomes and exposomes, HairDB, was introduced in Chapter 2. The database systematically compiles all reported hair chemicals through text-mining and manually incorporated chemicals with a high likelihood of being deposited on the surface of hair. HairDB contains 4191 unique chemicals from 9214 articles, 172 of which were further categorized as biomarkers. The user-friendly web interface of HairDB will facilitate the widespread adoption of hair as a matrix for research. Chapter 3 presents a novel bioinformatic pipeline for global-scale hair exposome studies, including optimized feature acquisition, feature quality assurance, intensity correction, top-down annotation, and molecular formula prediction using BUDDY. Out of the 26119 features detected, 316 were annotated using NIST20 and MS-DIAL spectral libraries. BUDDY predicted molecular formulas for 4533 unidentified metabolic features, with 4270 being successfully predicted. Global optimization of chemical annotation was applied to detect 1755 potential transformations, with 579 being between identified metabolic features and those with predicted molecular formulas through known chemical reactions. Using HairDB, 43 unique metabolites were found with corresponding literature. The predicted molecular formulas were also used to search HairDB, resulting in 275 hits. The development of HairDB and the annotation pipeline for hair metabolome offer valuable resources for researchers in the fields of metabolomics and exposomics.
View record
Development of mass spectrometry-based untargeted metabolomics for precision health (2022)
As the most recently emerged “omics”, metabolomics grabbed attention in human health studies by measuring thousands of small-molecule metabolites in a wide range of biological samples. As the downstream products in the biological pathway, metabolites are regarded as the closest link to the phenotypes. Small stimuli in the human body will cause relatively huge changes in the level of metabolites. Liquid chromatography-mass spectrometry (LC-MS) is the mainstay in metabolomics research due to its high throughput, sensitivity, and reliable analysis of metabolites. Nevertheless, two of the main challenges in LC-MS based metabolomics are 1) how to apply metabolomics in studying human health and 2) apart from commonly used biological samples, including serum, plasma, and urine, how to develop a methodology of new biological samples that can be adapted to specific human health research. To address those challenges, in Chapter 2, I integrated metabolomics with metagenomics to examine human gut health. 13-species metagenomic signature was selected by random forest machine learning and achieved high diagnostic accuracy in differentiating hepatic decompensation in NAFLD-related cirrhosis. The signature was cross-validated by metabolomics. 32 metabolites and 15 metabolites from serum and feces, respectively, were found to be significantly linked to 13-discriminatory species, suggesting that the identified discriminatory species may play important roles in the progression from compensated to decompensated cirrhosis. This multi-omics study yields new avenues for identifying novel targets for therapy and microbial biomarkers of hepatic decompensation, a worldwide human disease. In Chapter 3, I integrated plasma metabolomics and proteomics to examine the health conditions of highly trained females and males following acute, severe-intensity exercise. Metabolomic and proteomic homeostasis were substantially perturbed. Through statistical analysis, some metabolites and proteins were found to be closely linked to high-intensity exercise. This multi-omics study was a powerful tool to study molecular responses to acute exercise and provided a new insight to exercise-bolstered human health. In Chapter 4, I developed a new methodology to track skin secretion. Our high-performance workflow was readily applied to a wide range of skin metabolomics research to gain a better understanding of the molecular signatures on skin that link to human health and disease.
View record
The comprehensive analysis of different MS/MS acquisition modes and mass spectrometry related applications (2022)
For this work, two areas of metabolomics were investigated relating to the fundamentals of the field and application to different experiments. The first chapter was an assessment and comparison of the MSMS spectra generated from different acquisition modes. The chosen acquisition modes were data-dependant acquisition (DDA), data-independent acquisition (DIA), and enhanced insource-fragmentation (eISF) at a range of collision energies. The data was obtained by performing untargeted metabolomics on a urine sample and a standard mixture solution through a LC-MS platform while also covering multiple ionization modes. The spectra from the three modes were compared against each other through several factors that relate to the various ways MSMS spectra are used in a metabolomics workflow. These comparisons involved investigating the spectral purity, quality of reference matching results, structural similarity, and de novo annotation performance. It was found that DDA performed the best with eISF and DIA following. It was seen that eISF performed on-par or slightly better than DIA at higher collision energies. This indicates that the collision energy used will have a notable impact on the performance of the mode. The second chapter involves the metabolomics of pancreatic cell samples. The purpose of this was to determine the metabolic profile between the control and treated groups. The control was regular cancer cells from the MiaPaCa2 cell line while the treated groups had specific genes knocked out. The investigation was performed to gain insight into which metabolic pathways the knocked-out genes were involved in. Using a LC-MS platform it was found that 12 metabolites showed significant intensity differences between the groups. A literature review of these compounds highlighted possible metabolomic pathways affected such as polyamine metabolism. The last chapter focuses on a lipidomics experiment that was performed on the bacteria Thermotoga maritima to investigate the lipid content of the bacterial membranes. The samples relating to each fraction were run through the same LC-MS platform as above. It was seen that there were three significantly different lipids apart from the fatty acid, phosphatidylethanolamine, and phosphatidylinositol lipid classes. These classes have all been shown to be involved in membrane stability and transport.
View record
If this is your researcher profile you can log in to the Faculty & Staff portal to update your details and provide recruitment preferences.
Membership Status
Program Affiliations
Academic Unit(s)