Title Information
Title
Advancing Attention Mechanisms in Multimodal Deep Learning Models
Type of Resource (primo)
dissertations
Name: Personal
Name Part
Golovanevsky, Michal
Role
Role Term: Text
creator
Name: Personal
Name Part
Singh, Ritambhara
Role
Role Term: Text
Advisor
Name: Personal
Name Part
Eickhoff, Carsten
Role
Role Term: Text
Reader
Name: Personal
Name Part
Sun, Chen
Role
Role Term: Text
Reader
Name: Corporate
Name Part
Brown University. Department of Computer Science
Role
Role Term: Text
sponsor
Origin Information
Copyright Date
2026
Physical Description
Extent
xii, 98 p.
digitalOrigin
born digital
Note: thesis
Thesis (Ph. D.)--Brown University, 2026
Genre (aat)
theses
Abstract
Multimodal models are central to modern Artificial Intelligence systems, with growing relevance in domains such as healthcare. Their success is often credited to attention mechanisms, which enable rich cross-modal interactions. However, attention scales poorly when more than two modalities are integrated, and its internal functions remain poorly understood. This thesis addresses the scalability and interpretability of attention in multimodal models by introducing novel integration mechanisms, interpretability pipelines, and targeted interventions. In the first chapter, I introduce One-Versus-Others attention, a scalable alternative to self- and cross-attention that reduces computational cost while preserving accuracy in high-modality settings such as clinical data integration. In the second chapter, I introduce NOTICE, a vision–language causal mediation framework that enables the discovery of human-interpretable functions in attention mechanisms, revealing distinct roles for cross- and self-attention. In the third chapter, I analyze the internal competition between memorized priors and new perceptual input in multimodal models, demonstrating that steering vectors can reallocate attention to favor either prior knowledge or new visual evidence. In the final chapter, I examined how specific groups of attention heads induce prompt-copying behavior in vision–language models, leading to systematic prompt-induced hallucinations on numerical and semantic tasks. Through targeted ablation, I showed that these hallucinations can be reduced without retraining, and that the mechanisms by which copying is suppressed differ across models while producing a shared shift toward visual grounding. The findings presented in this thesis deepen our understanding of attention in multimodal models, offering concrete tools for scalability and interpretability. More broadly, they establish foundations for building multimodal systems that are trustworthy and reliable in real-world applications.
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/01034365")
Topic
Natural language processing (Computer science)
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/00817247")
Topic
Artificial intelligence
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/02032663")
Topic
Deep learning (Machine learning)
Language
Language Term (ISO639-2B)
English
Record Information
Record Content Source (marcorg)
RPB
Record Creation Date (encoding="iso8601")
20260427