5 papers
Buffer replay enhances the robustness of multimodal learning under missing-modality
Hongye Zhu, Xuan Liu, Yanwen Ba +2
Missing modalities consistently lead to significant performance degradation in multimodal models. Existing approaches either synthesize missing modalities at high computational cos…
Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
Wenxiao Wu, Jing-Hao Xue, Chengming Xu +5
Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedde…
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
Jingwei Zhao, Yuhua Wen, Qifei Li +8
Intent recognition aims to identify users' underlying intentions, traditionally focusing on text in natural language processing. With growing demands for natural human-computer int…
UMBRAE: Unified Multimodal Brain Decoding
Weihao Xia, Raoul de Charette, Cengiz Ãztireli +1
We address prevailing challenges of the brain-powered research, departing from the observation that the literature hardly recover accurate spatial information and require subject-s…
DREAM: Visual Decoding from Reversing Human Visual System
Weihao Xia, Raoul de Charette, Cengiz Ãztireli +1
In this work we present DREAM, an fMRI-to-image method for reconstructing viewed images from brain activities, grounded on fundamental knowledge of the human visual system. We craf…