From the 1 of 9 linked papers with an AI index.
9 papers
CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting
Jiayu Ding, Meilu Song, Yun Chen +2
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit inten…
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2
The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing
Feng Wang, Canmiao Fu, Zhipeng Huang +3
Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all histor…
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
Kuanwei Lin, Wenhao Zhang, Ge Li
Video large multimodal models increasingly face a scalability bottleneck: long videos produce excessively long visual-token sequences, which sharply increase memory and latency dur…
3D Instruction Ambiguity Detection
Jiayu Ding, Haoran Tang, Hongbo Jin +2
In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like "Pass me the vial" in a surgical setting could lead to catastrophic errors. Yet,…
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
Hongbo Jin, Jiayu Ding, Siyi Xie +2
Recent advancements in Large Language Models have successfully transitioned towards System 2 reasoning, yet applying these paradigms to video understanding remains challenging. Whi…