4 papers
DLAM: Distributional Latent Actions with Temporal Constraints
Zuojin Tang, Feifan Luo, Haoyun Liu +10
The paper introduces DLAM, a distributional latent-action model that encodes video transitions as diagonal Gaussians with temporal constraints, improving reconstruction consistency…
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
Zuojin Tang, Haoyun Liu, Xinyuan Chang +11
Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world…
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy
Zuojin Tang, Shengchao Yuan, Xiaoxin Bai +4
Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameterized on top of a pretrained VLA…
VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making
Zuojin Tang, Bin Hu, Chenyang Zhao +3
Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input si…