- Title Information
- Title
- Statistical Framework for Handling Variable Importance and Missing Data in Electronic Health Records
- Type of Resource (primo)
- dissertations
- Name:
Personal
- Name Part
- Lewis, Nickolas
- Role
- Role Term:
Text
- creator
- Name:
Personal
- Name Part
- Hogan, Joseph
- Role
- Role Term:
Text
- Advisor
- Name:
Personal
- Name Part
- Oganisian, Arman
- Role
- Role Term:
Text
- Reader
- Name:
Personal
- Name Part
- Steingrimsson, Jon
- Role
- Role Term:
Text
- Reader
- Name:
Personal
- Name Part
- Kantor, Rami
- Role
- Role Term:
Text
- Reader
- Name:
Corporate
- Name Part
- Brown University. Department of Biostatistics
- Role
- Role Term:
Text
- sponsor
- Origin Information
- Copyright Date
- 2026
- Physical Description
- Extent
- xix, 121 p.
- digitalOrigin
- born digital
- Note:
thesis
- Thesis (Ph. D.)--Brown University, 2026
- Genre (aat)
- theses
- Abstract
- Consistent medical care is essential for the health of people living with HIV (PLWH). Regular engagement in care improves access to antiretroviral therapy (ART), reduces progression to Acquired Immune Deficiency Syndrome (AIDS), and is associated with improved survival rates compared to PLWH who do not receive regular medical care. Population-level engagement in HIV care is commonly summarized through the HIV care cascade, a conceptual framework that defines key benchmarks for monitoring the effectiveness of HIV healthcare systems and a guideline for identifying gaps in care. The objective of this dissertation is to develop and implement data-driven tools to predict retention in HIV care to support clinical decision making at the Academic Model Providing Access to Healthcare (AMPATH), a large network of clinics providing HIV care in western Kenya. Currently, AMPATH dedicates considerable resources to patient outreach following a missed visit; however, the human resources used for outreach are not unlimited. Advanced outreach conducted prior to a scheduled return visit may increase the likelihood of continued engagement in care, motivating the development of models that predict a patient’s risk of missing an upcoming scheduled clinic visit. The first aim develops a principled statistical framework for predicting visit-level disengagement that explicitly aligns both model formulation and sampling strategies with the underlying data-generating process at each scheduled return visit. This framework uses a flexible Bayesian multinomial model that accounts for competing risks in the outcome space while simultaneously addressing missingness in the covariate space. The second aim focuses on interpretability by developing methods for local variable importance to support personalized outreach. We introduce a sampling-based approach grounded in predictive mean matching that samples from appropriate conditional distributions. We conduct simulations to demonstrate that this approach is scalable, robust to model misspecification, and applicable across a wide range of covariate types. The third aim examines how missing data at both model development and deployment affects global variable importance estimation. We extend a model-agnostic approach, Leave Out Covariates (LOCO), to explicitly account for missingness and use extensive simulations to evaluate its properties.
- Subject
- Topic
- HIV/AIDS
- Subject
- Topic
- missing data
- Subject
- Topic
- Bayesian Machine Learning
- Subject
- Topic
- Variable Importance
- Subject
- Topic
- Electronic Health Records
- Language
- Language Term (ISO639-2B)
- English
- Record Information
- Record Content Source (marcorg)
- RPB
- Record Creation Date
(encoding="iso8601")
- 20260427