7 papers
DLAM: Distributional Latent Actions with Temporal Constraints
Zuojin Tang, Feifan Luo, Haoyun Liu +10
The paper introduces DLAM, a distributional latent-action model that encodes video transitions as diagonal Gaussians with temporal constraints, improving reconstruction consistency…
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Ronghan Chen, Yandan Yang, Zuojin Tang +18
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack expl…
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement
Qianhan Feng, Wenshuo Li, Tong Lin +1
Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks. Meanwhile, the great need for capable aritificial intelligence on mobile…
Long-Text-to-Image Generation via Compositional Prompt Decomposition
Jen-Yuan Huang, Tong Lin, Yilun Du
While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are descriptive paragraphs. This…
ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning
Yandan Yang, Shuang Zeng, Tong Lin +11
Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one-brain, many-forms'' paradigm. Progress is hinder…
FARTrack: Fast Autoregressive Visual Tracking with High Performance
Guijie Wang, Tong Lin, Yifan Bai +4
Inference speed and tracking performance are two critical evaluation metrics in the field of visual tracking. However, high-performance trackers often suffer from slow processing s…