The recent proliferation of Multimodal Large Language Models (MLLMs) - particularly vision-language frameworks such as the LLaVA (Large Language-and-Vision Assistant) family - has heralded a new era of artificial intelligence characterized by sophisticated reasoning and instruction-following capabilities. However, this profound expressive power is inextricably linked to a significant computational bottleneck. In standard architectural formulations, MLLMs ingest the entirety of the visual tokens extracted by a pre-trained vision encoder. Because the self-attention mechanism at the core of the underlying Large Language Models (LLMs) scales quadratically with sequence length, processing an uncompressed, high-resolution grid of visual tokens becomes profoundly resource-intensive. To address this computational inefficiency, a dedicated module is proposed to systematically aggregate the vast set of visual features extracted by a vision encoder into a highly compact, yet semantically rich, set of dense vectors. By acting as an information bottleneck, this module aims to preserve the foundational semantic integrity required for complex multimodal reasoning while drastically reducing the quadratic computational overhead of LLMs. To determine the most effective compression strategy, three distinct aggregation paradigms are investigated. First, text-agnostic vision-only aggregation operates exclusively on the visual modality, compressing features through self-supervised reconstruction without relying on textual context. Second, text-influenced vision-centric aggregation incorporates textual influence directly into the training objective, conditioning the model to select task-relevant visual features. Finally, text-guided vision aggregation introduces a dynamic, context-aware mechanism where textual features actively govern the visual aggregation process. Through a rigorous evaluation of these paradigms across a comprehensive suite of downstream tasks, this research seeks to establish a scalable methodology for visual information synthesis, proving that the majority of visual tokens are not necessary, but rather a small set of dense ones are enough to achieve good results in downstream tasks, overcoming in thereby the computational bottlenecks of current MLLMs.

Less Vision Is All You Need: Reducing Visual Context in MLLMs

RICCIARDI, NICOLA
2025/2026

Abstract

The recent proliferation of Multimodal Large Language Models (MLLMs) - particularly vision-language frameworks such as the LLaVA (Large Language-and-Vision Assistant) family - has heralded a new era of artificial intelligence characterized by sophisticated reasoning and instruction-following capabilities. However, this profound expressive power is inextricably linked to a significant computational bottleneck. In standard architectural formulations, MLLMs ingest the entirety of the visual tokens extracted by a pre-trained vision encoder. Because the self-attention mechanism at the core of the underlying Large Language Models (LLMs) scales quadratically with sequence length, processing an uncompressed, high-resolution grid of visual tokens becomes profoundly resource-intensive. To address this computational inefficiency, a dedicated module is proposed to systematically aggregate the vast set of visual features extracted by a vision encoder into a highly compact, yet semantically rich, set of dense vectors. By acting as an information bottleneck, this module aims to preserve the foundational semantic integrity required for complex multimodal reasoning while drastically reducing the quadratic computational overhead of LLMs. To determine the most effective compression strategy, three distinct aggregation paradigms are investigated. First, text-agnostic vision-only aggregation operates exclusively on the visual modality, compressing features through self-supervised reconstruction without relying on textual context. Second, text-influenced vision-centric aggregation incorporates textual influence directly into the training objective, conditioning the model to select task-relevant visual features. Finally, text-guided vision aggregation introduces a dynamic, context-aware mechanism where textual features actively govern the visual aggregation process. Through a rigorous evaluation of these paradigms across a comprehensive suite of downstream tasks, this research seeks to establish a scalable methodology for visual information synthesis, proving that the majority of visual tokens are not necessary, but rather a small set of dense ones are enough to achieve good results in downstream tasks, overcoming in thereby the computational bottlenecks of current MLLMs.
2025
Multimodal LLMs
Token Compression
Efficiency
Text-Guided
Multimodal Reasoning
File in questo prodotto:
File Dimensione Formato  
Ricciardi.Nicola.pdf

accesso aperto

Dimensione 12.33 MB
Formato Adobe PDF
12.33 MB Adobe PDF Visualizza/Apri

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14251/7268