4 papers
World2Act: Latent Action Post-Training from World Model Dynamics
An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6
World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…
ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning
Tuan Van Vo, Tan Q. Nguyen, Khang Nguyen +5
Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructi…
Counting to Four is still a Chore for VLMs
Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo
Vision--language models (VLMs) have achieved impressive performance on complex multimodal reasoning tasks, yet they still fail on simple grounding skills such as object counting. E…
ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
Tuan Van Vo, Tan Quang Nguyen, Khang Minh Nguyen +2
Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observations with linguistic instructi…