4 papers
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Ying Chen, Weizhen Li, Zhe Hu +7
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state…
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents
Ying Chen, Lihuang Fang, Rui Jiang +4
Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call terminal commitment. Behaviora…
RL: Reasoning 3D Layouts from Relative Spatial Relations
Zhifeng Gu, Yuqi Wang, Bing Wang
Relative spatial relations provide a compact representation of spatial structure and are fundamental to relative spatial reasoning in 3D layout generation. Recent works leverage Mu…
MMOne: Representing Multiple Modalities in One Scene
Zhifeng Gu, Bing Wang
Humans perceive the world through multimodal cues to understand and interact with the environment. Learning a scene representation for multiple modalities enhances comprehension of…