<mods:mods xmlns:mods="http://www.loc.gov/mods/v3" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-7.xsd"><mods:titleInfo><mods:title>Interrogating Cellular Heterogeneity with Traditional Machine Learning and Single-Cell Foundation Models</mods:title></mods:titleInfo><mods:typeOfResource authority="primo">dissertations</mods:typeOfResource><mods:name type="personal"><mods:namePart>Denadel, Alan</mods:namePart><mods:role><mods:roleTerm type="text">creator</mods:roleTerm></mods:role></mods:name><mods:name type="personal"><mods:namePart>Singh, Ritambhara</mods:namePart><mods:role><mods:roleTerm type="text">Reader</mods:roleTerm></mods:role></mods:name><mods:name type="personal"><mods:namePart>De Vito, Roberta</mods:namePart><mods:role><mods:roleTerm type="text">Reader</mods:roleTerm></mods:role></mods:name><mods:name type="personal"><mods:namePart>Crawford, Lorin</mods:namePart><mods:role><mods:roleTerm type="text">Advisor</mods:roleTerm></mods:role></mods:name><mods:name type="corporate"><mods:namePart>Brown University. Center for Computational Molecular Biology</mods:namePart><mods:role><mods:roleTerm type="text">sponsor</mods:roleTerm></mods:role></mods:name><mods:originInfo><mods:copyrightDate>2025</mods:copyrightDate></mods:originInfo><mods:physicalDescription><mods:extent>21, 210 p.</mods:extent><mods:digitalOrigin>born digital</mods:digitalOrigin></mods:physicalDescription><mods:note type="thesis">Thesis (Ph. D.)--Brown University, 2025</mods:note><mods:genre authority="aat">theses</mods:genre><mods:abstract>Over the last decade, there have been dramatic improvements in assays that can
measure RNA in single cells, collectively known as scRNA-seq. These methods have
revealed unexpected heterogeneity in populations of cells that were previously thought
to be homogeneous and have enabled researchers to study how genes are co-expressed in
various cell types and cell states, across both healthy and diseased cells. Understanding the
complex dynamics of how cells of various types are impacted by disease has the potential
to allow us to develop targeted therapies, but there are methodological challenges to doing
so accurately. It is important for computational tools to guide researchers down useful
research paths, rather than dead ends. There are now thousands of computational tools for
analyzing scRNA-seq data and so-called single-cell foundation models have been recently
developed that aim to pre-train on prior studies with the goal of improved performance
when fine-tuned on multiple downstream tasks. The statistical analysis of scRNA-seq
data often involves generating and testing hypotheses using the same data, known as
“double-dipping”, which produces highly inflated P -values and can lead to false discoveries.
This dissertation presents three contributions to the fast-moving field of scRNA-seq that
demonstrate its utility in biological discovery, solve a key computational step in the
analysis of scRNA-seq data, and evaluates the recently developed single-cell foundation
models. To be specific, this dissertation (1) shows the importance of measuring cell
state in pancreatic ductal adenocarcinoma and identifies cell states that modulate drug
resistance, (2) develops an algorithm for detecting distinct cell types and cell states in an
unsupervised fashion (while preventing over-clustering) by controlling for the phenomenon
of data “double-dipping”, and (3) evaluates the role of pre-training dataset size and
diversity on single-cell foundation model performance.analysis of scRNA-seq data,
and evaluates the recently developed single-cell foundation models. To be specific, this
dissertation (1) shows the importance of measuring cell state in pancreatic ductal
adenocarcinoma and identifies cell states that modulate drug resistance, (2) develops
an algorithm for detecting distinct cell types and cell states in an unsupervised fashion
(while preventing over-clustering) by controlling for the phenomenon of data “double-dipping”,
and (3) evaluates the role of pre-training dataset size and diversity on single-cell foundation
model performance.</mods:abstract><mods:subject><mods:topic>RNA sequencing</mods:topic></mods:subject><mods:subject><mods:topic>unsupervised clustering</mods:topic></mods:subject><mods:subject><mods:topic>pancreatic cancer</mods:topic></mods:subject><mods:subject><mods:topic>knockoffs</mods:topic></mods:subject><mods:language><mods:languageTerm authority="iso639-2b">English</mods:languageTerm></mods:language><mods:recordInfo><mods:recordContentSource authority="marcorg">RPB</mods:recordContentSource><mods:recordCreationDate encoding="iso8601">20250707</mods:recordCreationDate></mods:recordInfo></mods:mods>