3 papers
cs.CV2026
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
Hao Vo, Khoa Vo, Phu Loc Nguyen +10
Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity acros…
cs.RO2025
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Khoa Vo, Taisei Hanyu, Yuki Ikebe +8
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action…
cs.RO2025
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Nhat Chung, Taisei Hanyu, Toan Nguyen +9
As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especiall…