collaborators

5 papers

cs.CV2026

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

Hao Vo, Phu Loc Nguyen, Khoa Vo +7

Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Auton…

cs.CV2026

TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

Duc Nguyen, Sieu Tran, Hao Vo +6

Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent video slot-attention methods propa…

cs.CV2026

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

Hao Vo, Khoa Vo, Phu Loc Nguyen +10

Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity acros…

cs.CV2026

Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning

Sieu Tran, Duc Nguyen, Hao Vo +2

Unsupervised video object-centric learning aims to decompose dynamic scenes into persistent, object-level representations without supervision. However, existing slot-based methods…

cs.RO2026

CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models

Khoa Vo, Sieu Tran, Taisei Hanyu +8

Vision-Language-Action (VLA) models promise generalist robot manipulation, but are typically trained and deployed as short-horizon policies that assume the latest observation is su…