4 papers · 1 filter
OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection
Xiaoyan Wei, Zhimin Yao, Ruilin Yang +4
Recent unified open-vocabulary detection (OVD) supports heterogeneous prompts, including text queries, visual exemplars, and their combinations, but often rely on increasingly comp…
Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
Zhongkuan Mao, Xianjie Liu, Tianyu Meng +9
High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect i…
Current World Models Lack a Persistent State Core
Jinpeng Lu, Dexu Zhu, Haoyuan Shi +8
World models are increasingly regarded as a decisive step toward artificial general intelligence, yet modeling the physical world demands more than rendering convincing frames on d…
EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models
Haozhe Shan, Xiancong Ren, Han Dong +9
While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on question-answering or multiple-…