Description
- Abstract:
- Over the last decade, there have been dramatic improvements in assays that can measure RNA in single cells, collectively known as scRNA-seq. These methods have revealed unexpected heterogeneity in populations of cells that were previously thought to be homogeneous and have enabled researchers to study how genes are co-expressed in various cell types and cell states, across both healthy and diseased cells. Understanding the complex dynamics of how cells of various types are impacted by disease has the potential to allow us to develop targeted therapies, but there are methodological challenges to doing so accurately. It is important for computational tools to guide researchers down useful research paths, rather than dead ends. There are now thousands of computational tools for analyzing scRNA-seq data and so-called single-cell foundation models have been recently developed that aim to pre-train on prior studies with the goal of improved performance when fine-tuned on multiple downstream tasks. The statistical analysis of scRNA-seq data often involves generating and testing hypotheses using the same data, known as “double-dipping”, which produces highly inflated P -values and can lead to false discoveries. This dissertation presents three contributions to the fast-moving field of scRNA-seq that demonstrate its utility in biological discovery, solve a key computational step in the analysis of scRNA-seq data, and evaluates the recently developed single-cell foundation models. To be specific, this dissertation (1) shows the importance of measuring cell state in pancreatic ductal adenocarcinoma and identifies cell states that modulate drug resistance, (2) develops an algorithm for detecting distinct cell types and cell states in an unsupervised fashion (while preventing over-clustering) by controlling for the phenomenon of data “double-dipping”, and (3) evaluates the role of pre-training dataset size and diversity on single-cell foundation model performance.analysis of scRNA-seq data, and evaluates the recently developed single-cell foundation models. To be specific, this dissertation (1) shows the importance of measuring cell state in pancreatic ductal adenocarcinoma and identifies cell states that modulate drug resistance, (2) develops an algorithm for detecting distinct cell types and cell states in an unsupervised fashion (while preventing over-clustering) by controlling for the phenomenon of data “double-dipping”, and (3) evaluates the role of pre-training dataset size and diversity on single-cell foundation model performance.
- Notes:
- Thesis (Ph. D.)--Brown University, 2025
Citation
Denadel, Alan,
"Interrogating Cellular Heterogeneity with Traditional Machine Learning and Single-Cell Foundation Models"
(2025).
Center for Computational Molecular Biology Theses and Dissertations.
Brown Digital Repository. Brown University Library.
https://repository.library.brown.edu/studio/item/bdr:uchnt3tj/
Relations
Collection:
-
Center for Computational Molecular Biology Theses and Dissertations
Theses and Dissertations for the Center for Computational Molecular Biology department....