1 paper
Daniel Carlström Schad, Shrey Dixit, Janis Keck +3
We present VIBE, a two-stage Transformer that fuses multi-modal video, audio, and text features to predict fMRI activity. Representations from open-source models (Qwen2.5, BEATs, W…