Brown University

Advancing Attention Mechanisms in Multimodal Deep Learning Models

Description

Abstract:
Multimodal models are central to modern Artificial Intelligence systems, with growing relevance in domains such as healthcare. Their success is often credited to attention mechanisms, which enable rich cross-modal interactions. However, attention scales poorly when more than two modalities are integrated, and its internal functions remain poorly understood. This thesis addresses the scalability and interpretability of attention in multimodal models by introducing novel integration mechanisms, interpretability pipelines, and targeted interventions. In the first chapter, I introduce One-Versus-Others attention, a scalable alternative to self- and cross-attention that reduces computational cost while preserving accuracy in high-modality settings such as clinical data integration. In the second chapter, I introduce NOTICE, a vision–language causal mediation framework that enables the discovery of human-interpretable functions in attention mechanisms, revealing distinct roles for cross- and self-attention. In the third chapter, I analyze the internal competition between memorized priors and new perceptual input in multimodal models, demonstrating that steering vectors can reallocate attention to favor either prior knowledge or new visual evidence. In the final chapter, I examined how specific groups of attention heads induce prompt-copying behavior in vision–language models, leading to systematic prompt-induced hallucinations on numerical and semantic tasks. Through targeted ablation, I showed that these hallucinations can be reduced without retraining, and that the mechanisms by which copying is suppressed differ across models while producing a shared shift toward visual grounding. The findings presented in this thesis deepen our understanding of attention in multimodal models, offering concrete tools for scalability and interpretability. More broadly, they establish foundations for building multimodal systems that are trustworthy and reliable in real-world applications.
Notes:
Thesis (Ph. D.)--Brown University, 2026

Citation

Golovanevsky, Michal, "Advancing Attention Mechanisms in Multimodal Deep Learning Models" (2026). Computer Science Theses and Dissertations. Brown Digital Repository. Brown University Library. https://repository.library.brown.edu/studio/item/bdr:cfhchg2e/

Relations

Collection: