From the 1 of 6 linked papers with an AI index.
6 papers
DLAM: Distributional Latent Actions with Temporal Constraints
Zuojin Tang, Feifan Luo, Haoyun Liu +10
The paper introduces DLAM, a distributional latent-action model that encodes video transitions as diagonal Gaussians with temporal constraints, improving reconstruction consistency…
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Ronghan Chen, Yandan Yang, Zuojin Tang +18
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack expl…
World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
Junjin Xiao, Yandan Yang, Xinyuan Chang +5
Vision-Language-Action (VLA) models trained via imitation learning suffer from significant performance degradation in data-scarce scenarios due to their reliance on large-scale dem…
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
Yandan Yang, Shuang Zeng, Tong Lin +11
Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one-brain, many-forms'' paradigm. Progress is hinder…
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
Yuzhi Chen, Ronghan Chen, Dongjie Huo +11
Video-based world models offer a powerful paradigm for embodied simulation and planning, yet state-of-the-art models often generate physically implausible manipulations - such as o…
Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models
Zhejia Cai, Yandan Yang, Xinyuan Chang +5
Latent Action Models (LAMs) enable Vision- Language-Action (VLA) systems to learn semantic action representations from large-scale unannotated data. Yet, we identify two bottleneck…