6 papers
WALL-SS: Scaling Long-horizon World Models via Next-Scale Autoregression
Maeve Zhang, Rain Sun, Xiang Wang +22
Generative world models provide robots with predictive models of how the world evolves under interaction, with growing potential for simulation, planning, policy evaluation, and ro…
HOST:Robots Acquire Manipulation Skills in Seconds from a Single Human Video
Guangyan Chen, Meiling Wang, Te Cui +9
The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-…
X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance
Rime Wen, Zehan Liu, Shawn Qin +4
Streaming text-to-speech is essential for low-latency spoken dialogue systems, yet many systems wait for sentence-level text and are therefore only pseudo-streaming. True token-lev…
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Miracle Kang, Lights Shi, Lucy Liang +10
Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions prim…
DMuon: Efficient Distributed Muon Training with Near-Adam Overhead
Vincent Chen, Starrick Liu, Regis Cheng +8
Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads. The matrix-awar…
Igniting VLMs toward the Embodied Space
Andy Zhai, Brae Liu, Bruno Fang +17
While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferrin…