7 papers
UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation
Zeyuan Yang, Hao-Wei Chen, Xueyang Yu +5
Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single architecture. While autoregressiv…
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Reuben Tan, Baolin Peng, Zhengyuan Yang +16
Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-bas…
Action Images: End-to-End Policy Learning via Multiview Video Generation
Haoyu Zhen, Zixian Gao, Qiao Sun +7
World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, exist…
Fast Spatial Memory with Elastic Test-Time Training
Ziqiao Ma, Xueyang Yu, Haoyu Zhen +3
Large Chunk Test-Time Training (LaCT) has shown strong performance on long-context 3D reconstruction, but its fully plastic inference-time updates remain vulnerable to catastrophic…
MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
Yuncong Yang, Jiageng Liu, Zheyuan Zhang +5
Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision-language mode…
Learning 3D Persistent Embodied World Models
Siyuan Zhou, Yilun Du, Yuncong Yang +4
The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions a…