52 papers
Attentive multilayer fusion for vision transformers
Laure Ciernik, Marco Morik, Lukas Thede +4
The paper introduces Attentive Layer Fusion (ALF), a method that dynamically combines representations from all layers of a Vision Transformer to improve linear probing on downstrea…
TimeSAE: Causal Sparse Decoding for Faithful Explanations of Black-Box Time Series Models
Khalid Oublal, Quentin Bouniot, Qi Gan +2
As black box models and pretrained models gain traction in time series applications, understanding and explaining their predictions becomes increasingly vital, especially in high-s…
From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
Sena Korkut, MarÃa Alejandra Bravo Sarmiento, Sanghwan Kim +1
High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers shoul…
SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
Simon Roschmann, Paul Krzakala, Sonia Mazelet +2
The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits thi…
Mantis: Lightweight Foundation Model for Time Series Classification
Vasilii Feofanov, Songkang Wen, Shifeng Xie +10
While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored, with existing literature predominantly f…
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
Sanghwan Kim, Rui Xiao, Stephan Alaniz +2
Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long v…