2 papers
cs.LG2025
VIBE: Video-Input Brain Encoder for fMRI Response Modeling
Daniel Carlström Schad, Shrey Dixit, Janis Keck +3
We present VIBE, a two-stage Transformer that fuses multi-modal video, audio, and text features to predict fMRI activity. Representations from open-source models (Qwen2.5, BEATs, W…
cs.LG2025
Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
Shrey Dixit, Kayson Fakhar, Fatemeh Hadaeghi +3
Neural networks now generate text, images, and speech with billions of parameters, producing a need to know how each neural unit contributes to these high-dimensional outputs. Exis…