7 papers
Cache: Accelerating World Action Models with Cross Inference Chunk Cache
Weisen Zhao, Lam Nguyen, Zhicong Lu +1
World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-modeling objective lets them lea…
Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
Kunyang Li, Mubarak Shah, Yuzhang Shang
Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attention leads to quadratic comput…
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
Zhiben Chen, Youpeng Zhao, Yang Sui +2
Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive (AR) models, offering better hardware utilization and bidirectional context thro…
Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls
Abdul Mohaimen Al Radi, Kunyang Li, Yuzhang Shang +2
Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural language rather than low-lev…
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Yu Shang, Yinzhou Tang, Yiding Ma +22
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…
Good Agentic Friends Do Not Just Give Verbal Advice: They Can Update Your Weights
Wenrui Bao, Huan Wang, Jian Wang +3
Multi-agent LLM systems usually collaborate by exchanging natural-language messages. This interface is simple and interpretable, but it forces each sender's intermediate computatio…