Title Information
Title
Interrogating Cellular Heterogeneity with Traditional Machine Learning and Single-Cell Foundation Models
Type of Resource (primo)
dissertations
Name: Personal
Name Part
Denadel, Alan
Role
Role Term: Text
creator
Name: Personal
Name Part
Singh, Ritambhara
Role
Role Term: Text
Reader
Name: Personal
Name Part
De Vito, Roberta
Role
Role Term: Text
Reader
Name: Personal
Name Part
Crawford, Lorin
Role
Role Term: Text
Advisor
Name: Corporate
Name Part
Brown University. Center for Computational Molecular Biology
Role
Role Term: Text
sponsor
Origin Information
Copyright Date
2025
Physical Description
Extent
21, 210 p.
digitalOrigin
born digital
Note: thesis
Thesis (Ph. D.)--Brown University, 2025
Genre (aat)
theses
Abstract
Over the last decade, there have been dramatic improvements in assays that can measure RNA in single cells, collectively known as scRNA-seq. These methods have revealed unexpected heterogeneity in populations of cells that were previously thought to be homogeneous and have enabled researchers to study how genes are co-expressed in various cell types and cell states, across both healthy and diseased cells. Understanding the complex dynamics of how cells of various types are impacted by disease has the potential to allow us to develop targeted therapies, but there are methodological challenges to doing so accurately. It is important for computational tools to guide researchers down useful research paths, rather than dead ends. There are now thousands of computational tools for analyzing scRNA-seq data and so-called single-cell foundation models have been recently developed that aim to pre-train on prior studies with the goal of improved performance when fine-tuned on multiple downstream tasks. The statistical analysis of scRNA-seq data often involves generating and testing hypotheses using the same data, known as “double-dipping”, which produces highly inflated P -values and can lead to false discoveries. This dissertation presents three contributions to the fast-moving field of scRNA-seq that demonstrate its utility in biological discovery, solve a key computational step in the analysis of scRNA-seq data, and evaluates the recently developed single-cell foundation models. To be specific, this dissertation (1) shows the importance of measuring cell state in pancreatic ductal adenocarcinoma and identifies cell states that modulate drug resistance, (2) develops an algorithm for detecting distinct cell types and cell states in an unsupervised fashion (while preventing over-clustering) by controlling for the phenomenon of data “double-dipping”, and (3) evaluates the role of pre-training dataset size and diversity on single-cell foundation model performance.analysis of scRNA-seq data, and evaluates the recently developed single-cell foundation models. To be specific, this dissertation (1) shows the importance of measuring cell state in pancreatic ductal adenocarcinoma and identifies cell states that modulate drug resistance, (2) develops an algorithm for detecting distinct cell types and cell states in an unsupervised fashion (while preventing over-clustering) by controlling for the phenomenon of data “double-dipping”, and (3) evaluates the role of pre-training dataset size and diversity on single-cell foundation model performance.
Subject
Topic
RNA sequencing
Subject
Topic
unsupervised clustering
Subject
Topic
pancreatic cancer
Subject
Topic
knockoffs
Language
Language Term (ISO639-2B)
English
Record Information
Record Content Source (marcorg)
RPB
Record Creation Date (encoding="iso8601")
20250707