Representation alignment has recently shown promise as a strategy to enhance Multimodal Large Language Models (MLLMs), by encouraging their internal activations to follow the representational geometry of strong external vision encoders. Existing approaches, however, usually impose this constraint on a predefined layer of the language backbone, neglecting the finer organization induced by Transformer attention heads. In this thesis, we introduce Head-Wise Representation Alignment (HeRA), a training strategy that brings cross-modal alignment to the level of individual attention heads. Grounded in the Platonic Representation Hypothesis, our method targets the preservation of the topological structure of visual representations, namely their local neighborhood relations, across modalities. Building on the Mutual K-Nearest Neighbor (MKNN) alignment metric, we design a contrastive objective that provides a differentiable surrogate for matching such local structures. During multimodal training, HeRA applies this objective to selected attention heads of the LLM, chosen according to their MKNN-based alignment score. Surprisingly, we observe that regularizing the least aligned heads leads to the strongest improvements. Experiments on multiple MLLMs and 18 benchmarks show that \ours consistently enhances performance on demanding vision-centric tasks, while also reducing visual hallucinations by limiting the model's tendency to rely excessively on linguistic priors.
Mind the Heads: Topological Representation Alignment for Multimodal LLMs
MELIS, FEDERICO
2025/2026
Abstract
Representation alignment has recently shown promise as a strategy to enhance Multimodal Large Language Models (MLLMs), by encouraging their internal activations to follow the representational geometry of strong external vision encoders. Existing approaches, however, usually impose this constraint on a predefined layer of the language backbone, neglecting the finer organization induced by Transformer attention heads. In this thesis, we introduce Head-Wise Representation Alignment (HeRA), a training strategy that brings cross-modal alignment to the level of individual attention heads. Grounded in the Platonic Representation Hypothesis, our method targets the preservation of the topological structure of visual representations, namely their local neighborhood relations, across modalities. Building on the Mutual K-Nearest Neighbor (MKNN) alignment metric, we design a contrastive objective that provides a differentiable surrogate for matching such local structures. During multimodal training, HeRA applies this objective to selected attention heads of the LLM, chosen according to their MKNN-based alignment score. Surprisingly, we observe that regularizing the least aligned heads leads to the strongest improvements. Experiments on multiple MLLMs and 18 benchmarks show that \ours consistently enhances performance on demanding vision-centric tasks, while also reducing visual hallucinations by limiting the model's tendency to rely excessively on linguistic priors.| File | Dimensione | Formato | |
|---|---|---|---|
|
Melis.Federico.pdf
accesso aperto
Dimensione
3.23 MB
Formato
Adobe PDF
|
3.23 MB | Adobe PDF | Visualizza/Apri |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14251/7265