6 papers
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
Filipe Laitenberger, Dawid Kopiczko, Cees G. M. Snoek +1
We introduce GateSkip, a simple residual-stream gating mechanism that enables token-wise layer skipping in decoder-only LMs. Each Attention/MLP branch is equipped with a sigmoid-li…
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
Mohammadreza Salehi, Shashanka Venkataramanan, Ioana Simion +3
Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of…
KV Cache Steering for Controlling Frozen LLMs
Max Belitsky, Dawid J. Kopiczko, Michael Dorkenwald +4
We propose cache steering, a lightweight method for implicit steering of language models via a one-shot intervention applied directly to the key-value cache. To validate its effect…
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
Mohammadreza Salehi, Nikolaos Apostolikas, Efstratios Gavves +2
In the realm of novelty detection, accurately identifying outliers in data without specific class information poses a significant challenge. While current methods excel in single-o…
Lost in Time: A New Temporal Benchmark for VideoLLMs
Daniel Cores, Michael Dorkenwald, Manuel Mucientes +2
Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating video models presents its o…
Better Language Models Exhibit Higher Visual Alignment
Jona Ruthardt, Gertjan J. Burghouts, Serge Belongie +1
How well do text-only large language models (LLMs) align with the visual world? We present a systematic evaluation of this question by incorporating frozen representations of vario…