8 papers
NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding
Sijin Yu, Zijiao Chen, Zhenyu Yang +7
Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically faithful surface-based models. We argue this is partly driven by in…
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
Yuanzhi Xu, Qian Gao, Jun Fan +4
The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to…
Efficient Agent: Optimizing Planning Capability for Multimodal Retrieval Augmented Generation
Yuechen Wang, Yuming Qiao, Dan Meng +4
Multimodal Retrieval-Augmented Generation (mRAG) has emerged as a promising solution to address the temporal limitations of Multimodal Large Language Models (MLLMs) in real-world s…
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
Jian Ma, Qirong Peng, Xu Guo +3
Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in…
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
Qi Wu, Quanlong Zheng, Yanhao Zhang +8
With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating…
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Zhenyi Liao, Qingsong Xie, Yanhao Zhang +4
Increasing attention has been placed on improving the reasoning capacities of multi-modal large language models (MLLMs). As the cornerstone for AI agents that function in the physi…