4 papers
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
Khoa Vo, Sieu Tran, Taisei Hanyu +8
Vision-Language-Action (VLA) models promise generalist robot manipulation, but are typically trained and deployed as short-horizon policies that assume the latest observation is su…
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
Taisei Hanyu, Nhat Chung, Huy Le +10
Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation representations can form a foundation for…
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Khoa Vo, Taisei Hanyu, Yuki Ikebe +8
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action…
GazeSearch: Radiology Findings Search Benchmark
Trong Thang Pham, Tien-Phat Nguyen, Yuki Ikebe +5
Medical eye-tracking data is an important information source for understanding how radiologists visually interpret medical images. This information not only improves the accuracy o…