2 papers
cs.CV2026
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
Riccardo Renzulli, Gabriele Spadaro, Shruthi Gowda +2
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated by the large number of visual tokens fed t…
cs.LG2026
Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs
Van-Tuan Tran, Shruthi Gowda, Merim Dzaferagic +1
Continual adaptation is essential for multimodal large language models (MLLMs) deployed across evolving domains, but the state-of-the-art MR-LoRA method highly relies on the assump…