<mods:mods xmlns:mods="http://www.loc.gov/mods/v3" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.loc.gov/mods/v3 http://www.loc.gov/standards/mods/v3/mods-3-7.xsd"><mods:titleInfo><mods:title>Advancing Attention Mechanisms in Multimodal Deep Learning Models</mods:title></mods:titleInfo><mods:typeOfResource authority="primo">dissertations</mods:typeOfResource><mods:name type="personal"><mods:namePart>Golovanevsky, Michal</mods:namePart><mods:role><mods:roleTerm type="text">creator</mods:roleTerm></mods:role></mods:name><mods:name type="personal"><mods:namePart>Singh, Ritambhara</mods:namePart><mods:role><mods:roleTerm type="text">Advisor</mods:roleTerm></mods:role></mods:name><mods:name type="personal"><mods:namePart>Eickhoff, Carsten</mods:namePart><mods:role><mods:roleTerm type="text">Reader</mods:roleTerm></mods:role></mods:name><mods:name type="personal"><mods:namePart>Sun, Chen</mods:namePart><mods:role><mods:roleTerm type="text">Reader</mods:roleTerm></mods:role></mods:name><mods:name type="corporate"><mods:namePart>Brown University. Department of Computer Science</mods:namePart><mods:role><mods:roleTerm type="text">sponsor</mods:roleTerm></mods:role></mods:name><mods:originInfo><mods:copyrightDate>2026</mods:copyrightDate></mods:originInfo><mods:physicalDescription><mods:extent>xii, 98 p.</mods:extent><mods:digitalOrigin>born digital</mods:digitalOrigin></mods:physicalDescription><mods:note type="thesis">Thesis (Ph. D.)--Brown University, 2026</mods:note><mods:genre authority="aat">theses</mods:genre><mods:abstract>Multimodal models are central to modern Artificial Intelligence systems, with growing relevance in domains
such as healthcare. Their success is often credited to attention mechanisms, which enable rich cross-modal
interactions. However, attention scales poorly when more than two modalities are integrated, and its internal
functions remain poorly understood. This thesis addresses the scalability and interpretability of attention
in multimodal models by introducing novel integration mechanisms, interpretability pipelines, and targeted
interventions. In the first chapter, I introduce One-Versus-Others attention, a scalable alternative to self- and
cross-attention that reduces computational cost while preserving accuracy in high-modality settings such as
clinical data integration. In the second chapter, I introduce NOTICE, a vision–language causal mediation
framework that enables the discovery of human-interpretable functions in attention mechanisms, revealing
distinct roles for cross- and self-attention. In the third chapter, I analyze the internal competition between
memorized priors and new perceptual input in multimodal models, demonstrating that steering vectors can
reallocate attention to favor either prior knowledge or new visual evidence. In the final chapter, I examined
how specific groups of attention heads induce prompt-copying behavior in vision–language models, leading
to systematic prompt-induced hallucinations on numerical and semantic tasks. Through targeted ablation, I
showed that these hallucinations can be reduced without retraining, and that the mechanisms by which copying
is suppressed differ across models while producing a shared shift toward visual grounding. The findings
presented in this thesis deepen our understanding of attention in multimodal models, offering concrete tools
for scalability and interpretability. More broadly, they establish foundations for building multimodal systems
that are trustworthy and reliable in real-world applications.</mods:abstract><mods:subject authority="fast" authorityURI="http://id.worldcat.org/fast" valueURI="http://id.worldcat.org/fast/01034365"><mods:topic>Natural language processing (Computer science)</mods:topic></mods:subject><mods:subject authority="fast" authorityURI="http://id.worldcat.org/fast" valueURI="http://id.worldcat.org/fast/00817247"><mods:topic>Artificial intelligence</mods:topic></mods:subject><mods:subject authority="fast" authorityURI="http://id.worldcat.org/fast" valueURI="http://id.worldcat.org/fast/02032663"><mods:topic>Deep learning (Machine learning)</mods:topic></mods:subject><mods:language><mods:languageTerm authority="iso639-2b">English</mods:languageTerm></mods:language><mods:recordInfo><mods:recordContentSource authority="marcorg">RPB</mods:recordContentSource><mods:recordCreationDate encoding="iso8601">20260427</mods:recordCreationDate></mods:recordInfo></mods:mods>