Description
- Abstract:
- This dissertation is motivated by complex data architectures—typical in genetics research—that break the assumptions of classical frequentist analyses. Predictors are not necessarily independent nor fewer in number than observations, and we utilize the linear mixed model paradigm to reflect collinearity through non-diagonal covariance matrices. Two sub-themes are present: (1) dissecting non-additive contributions to an outcome of interest; and (2) revealing latent, linear variation in an exploratory setting. The first and second chapters fall beneath the first sub-theme, albeit in different respects. Chapter 1 simultaneously addresses the enduring need for interpretability in machine learning and the difficulty of capturing the collective importance of multiple, related variables. GroupRATE is our novel measure of joint significance that extends existing univariate criterion RelATive cEntrality (RATE) to groups of inputs—for example, single nucleotide polymorphisms (SNPs) rolled up into genes—in a Bayesian neural network. While Chapter 1 considers the explanatory importance of raw features, Chapter 2 focuses on the explanatory importance of different types of genetic variation. Specifically, SNPs contribute both additive and pairwise interaction, or epistatic, effects to heritable phenotypes. The former is known as narrow-sense heritability, and the latter is a subset of broad-sense heritability, a catch-all descriptor for non-additive genetic variation. Genome-wide summary statistics that relate additive SNP effects to a phenotype have been shown to exhibit heterogeneity when conditioning on covariates such as genetic sex, but it is unknown how this phenomenon is reflected in narrow- and broad-sense heritability. We design an Epistatic and Additive Variance Estimation (EAVE) framework to quantify these respective sources of variation on a per-genomic region, per-cohort basis. The second sub-theme of this dissertation emerges in Chapter 3. We introduce Multi-Study Factor Analysis with Missingness (MSFAM), a factor analysis method to dissect common and study-specific latent variation across multiple studies when data are incomplete. This work is inspired by MSFA, which is operable only when data are fully observed. Although broadly applicable to any collection of related, incomplete data sets (surveys conducted at more than one site come to mind), one motivating use case is the discovery of genetic pathways from sparsely observed gene expression counts gathered from different patients.
- Notes:
- Thesis (Ph. D.)--Brown University, 2023
Citation
Udwin, Dana Lauren,
"Linear Mixed Models for Heterogeneity Estimation"
(2023).
Biostatistics Theses and Dissertations.
Brown Digital Repository. Brown University Library.
https://repository.library.brown.edu/studio/item/bdr:y7npnfry/
Relations
Collection:
-
Biostatistics Theses and Dissertations
Theses and Dissertations for the Biostatistics department....