17 papers
AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control
Yutian Cheng, Xiaojian Ma, Xianhao Wang +6
Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational ov…
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
Pengxiang Li, Zhi Gao, Bofei Zhang +8
Multimodal agents, which integrate a controller e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks.…
Chain Of Interaction Benchmark (COIN): When Reasoning meets Embodied Interaction
Xianhao Wang, Xiaojian Ma, Haozhe Hu +7
Generalist embodied agents must perform interactive, causally-dependent reasoning, continually interacting with the environment, acquiring information, and updating plans to solve…
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
Jiangyong Huang, Xiaojian Ma, Xiongkun Linghu +6
Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal. Despite recent progress, 3D VLMs still struggle with spatial reas…
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
Yapeng Mi, Yanpeng Zhao, Hengli Li +6
Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image gen…
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
Bofei Zhang, Zirui Shang, Zhi Gao +7
Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. Ho…