4 papers
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Miracle Kang, Lights Shi, Lucy Liang +10
Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions prim…
DMuon: Efficient Distributed Muon Training with Near-Adam Overhead
Vincent Chen, Starrick Liu, Regis Cheng +8
Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads. The matrix-awar…
WALL-WM: Carving World Action Modeling at the Event Joints
Shalfun Li, Victor Yao, Charles Yang +28
WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining, using semantically coherent…
Igniting VLMs toward the Embodied Space
Andy Zhai, Brae Liu, Bruno Fang +17
While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferrin…