5 papers · 1 filter
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
Fengyi Wu, Yifei Dong, Yilong Dai +7
We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentri…
RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
Xiyan Liu, Han Wang, Yuhu Wang +4
Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous…
Dynamic Double Space Tower
Weikai Sun, Shijie Song, Han Wang
The Visual Question Answering (VQA) task requires the simultaneous understanding of image content and question semantics. However, existing methods often have difficulty handling c…
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
Peiyuan Chen, Zecheng Zhang, Yiping Dong +2
Visual Question Answering (VQA) is a challenging task that requires systems to provide accurate answers to questions based on image content. Current VQA models struggle with comple…