From the 1 of 9 linked papers with an AI index.
9 papers
REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering
Caijun Yan, Yang Zhou, Meixing Shi +4
Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on ri…
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Fan Jiang, Zhaoxu Sun, Mengchao Wang +38
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA g…
Continual Video-MLLM Adaptation over Evolving Domains
Rui Cheng, Meixing Shi, Yuxiang Cai +3
Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remains underexplored. In real-wor…
Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction
Zengjie Chen, Yuxiang Cai, Jingcai Guo +3
Visual token reduction has emerged as an effective strategy for accelerating Multimodal Large Language Models (MLLMs). Many existing methods prune tokens by ranking text-visual att…
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
Yankai Jiang, Qiaoru Li, Binlu Xu +6
Recent research on medical MLLMs has gradually shifted its focus from image-level understanding to fine-grained, pixel-level comprehension. Although segmentation serves as the foun…