3 papers
cs.CV2026
Causally Debiased Latent Action Model for Embodied Action Conditioned World Models
Yufan Wei, Kun Zhou, Lingjun Mao +9
Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation,…
cs.AI2026
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
Yifei Shao, Kun Zhou, Ziming Xu +5
We study how to extend chain-of-thought (CoT) beyond language to better handle multimodal reasoning. While CoT helps LLMs and VLMs articulate intermediate steps, its text-only form…
cs.CV2025
Learning Plug-and-play Memory for Guiding Video Diffusion Models
Selena Song, Ziming Xu, Zijun Zhang +4
Diffusion Transformer(DiT) based video generation models have recently achieved impressive visual quality and temporal coherence, but they still frequently violate basic physical l…