Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
World2Act: Latent Action Post-Training from World Model Dynamics
An Dinh Vuong, Tuan Van Vo, Abdullah Sohail +6
World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene…
cs.CV2026
Counting to Four is still a Chore for VLMs
Duy Le Dinh Anh, Patrick Amadeus Irawan, Tuan Van Vo
Vision--language models (VLMs) have achieved impressive performance on complex multimodal reasoning tasks, yet they still fail on simple grounding skills such as object counting. E…
cs.CV2024
CathAction: A Benchmark for Endovascular Intervention Understanding
Baoru Huang, Tuan Vo, Chayun Kongtongvattana +35
Real-time visual feedback from catheterization analysis is crucial for enhancing surgical safety and efficiency during endovascular interventions. However, existing datasets are of…