4 papers
Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
Yufan Wei, Kun Zhou, Lingjun Mao +9
Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation,…
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
Hongyu Ding, Sizhuo Zhang, Ziming Xu +13
Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dom…
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
Yifei Shao, Kun Zhou, Ziming Xu +5
We study how to extend chain-of-thought (CoT) beyond language to better handle multimodal reasoning. While CoT helps LLMs and VLMs articulate intermediate steps, its text-only form…
Learning Plug-and-play Memory for Guiding Video Diffusion Models
Selena Song, Ziming Xu, Zijun Zhang +4
Diffusion Transformer(DiT) based video generation models have recently achieved impressive visual quality and temporal coherence, but they still frequently violate basic physical l…