10 papers
SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
Yuqing Feng, Jiawei Ma, Kevin Qinghong Lin +6
Surgical procedures unfold as structured and recurring clinical events, whose real-time understanding via intraoperative surgical videos is critical for intraoperative decision-mak…
OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs
Hao Vo, Phu Loc Nguyen, Khoa Vo +7
Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-world applications such as Auton…
DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
Hao Vo, Khoa Vo, Phu Loc Nguyen +10
Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity acros…
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
Khoa Vo, Sieu Tran, Taisei Hanyu +8
Vision-Language-Action (VLA) models promise generalist robot manipulation, but are typically trained and deployed as short-horizon policies that assume the latest observation is su…
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
Taisei Hanyu, Nhat Chung, Huy Le +10
Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation representations can form a foundation for…
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Khoa Vo, Taisei Hanyu, Yuki Ikebe +8
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action…