6 papers · 1 filter
HumanCLAW: Can Vision-Language Models Act Through a Body?
Li Siyao, Jiawei Gu, Shuai Liu +15
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task…
3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis
Yuhao Wang, Puyi Wang, Linjie Li +3
Most recent 3D reconstruction and editing systems operate on implicit and explicit representations such as NeRF, point clouds, or meshes. While these representations enable high-fi…
Gym-V: A Unified Vision Environment System for Agentic Vision Research
Fanqing Meng, Lingxiao Du, Jiawei Gu +9
As agentic systems increasingly rely on reinforcement learning from verifiable rewards, standardized ``gym'' infrastructure has become essential for rapid iteration, reproducibilit…
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
Jiawei Gu, Yunzhuo Hao, Huichen Will Wang +5
Multimodal reasoning requires iterative coordination between language and vision, yet it remains unclear what constitutes a meaningful interleaved chain of thought. We posit that t…
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
Zhaochen Su, Peng Xia, Hangyu Guo +12
Recent progress in multimodal reasoning has been significantly advanced by textual Chain-of-Thought (CoT), a paradigm where models conduct reasoning within language. This text-cent…
Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study
Jihai Zhang, Tianle Li, Linjie Li +2
Unified vision-language models (VLMs) aim to support both visual understanding and generation within a single framework, but it remains unclear when mixed training benefits both ca…