Brown University

Machine Learning Methods for Bias Correction and Precision Optimization Using Covariate Adjustment in Randomized Trials With Missing Data

Description

Abstract:
Adjusting for pre-specified baseline covariates in randomized trials can result in efficiency gains for the estimated effects. Even under misspecification of the adjustment model, the effect estimate remains unbiased under complete data. In missing data contexts, however, the misspecification of the adjustment model can result in biased treatment effect estimates. We investigate machine learning (ML) methods for the adjustment model and address two questions. Under complete data, we investigate whether adopting ML methods for the adjustment model enhances efficiency gains relative to a misspecified multiple linear regression model. For missing data, we explore whether ML adjustment improves efficiency while avoiding bias attributable to adjustment model misspecification. We demonstrate the potential of ML adjustment using simulated datasets of varied sample sizes ranging from 500 to 2000. Simulated datasets mimic a two-arm randomized trial evaluating a hypothetical treatment. We assume intention-to-treat analyses and estimate the treatment effect on a continuous primary outcome under complete, informative, and non-informative missing outcome data. All adjustment models, except the correct outcome model, are misspecified to represent uncertainty about the underlying outcome data-generating mechanism. The effect of differences in the proportion of variance explained by baseline covariates in the correct model is explored by varying the residual variance. Methods are illustrated via an application to a randomized trial. Under complete data, ML adjustment improves efficiency with gains directly related to the proportion of variance in the outcome explained by baseline covariates in the correct model. Relative to the unadjusted estimator, our best-performing ML algorithm (BART) improved efficiency by up to 36% and up to 7% compared to using a misspecified linear regression model when the proportion variance explained by the baseline covariates in the correct model is low, and more than double when high. Similar findings hold for missing data, with the degree of bias correction dependent on the missing data mechanism. Adopting ML for the adjustment model can result in efficiency gains even under model misspecification and missing data in randomized trials. Optimal efficiency improvements follow when the proportion of variance explained by covariates in the correct outcome model is high. However, caution should be exercised when the sample size is small. This research provides additional guidance for the appropriate use of ML covariate adjustment in randomized trials for efficiency gains.
Notes:
Thesis (Sc. M.)--Brown University, 2023

Citation

Okutse, Amos Ochieng, "Machine Learning Methods for Bias Correction and Precision Optimization Using Covariate Adjustment in Randomized Trials With Missing Data" (2023). Biostatistics Theses and Dissertations. Brown Digital Repository. Brown University Library. https://repository.library.brown.edu/studio/item/bdr:n56nec8a/

Relations

Collection: