2 papers
cs.SD2026
Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
Lukas Rauch, René Heinrich, Houtan Ghaffari +4
Although probing frozen models has become a standard evaluation paradigm, self-supervised learning in audio defaults to fine-tuning when pursuing state-of-the-art on AudioSet. A ke…
cs.CV2025
MIM-Refiner: A Contrastive Learning Boost from Intermediate Pre-Trained Representations
Benedikt Alkin, Lukas Miklautz, Sepp Hochreiter +1
We introduce MIM (Masked Image Modeling)-Refiner, a contrastive learning boost for pre-trained MIM models. MIM-Refiner is motivated by the insight that strong representations withi…