5 papers
SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
Taisei Hanyu, Nhat Chung, Huy Le +10
Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation representations can form a foundation for…
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Khoa Vo, Taisei Hanyu, Yuki Ikebe +8
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action…
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
Huy Le, Nhat Chung, Tung Kieu +2
Video Scene Graph Generation (VidSGG) aims to represent dynamic visual content by detecting objects and modeling their temporal interactions as structured graphs. Prior studies typ…
Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Nhat Chung, Taisei Hanyu, Toan Nguyen +9
As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especiall…
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
Huy Le, Nhat Chung, Tung Kieu +2
Text-video retrieval (TVR) systems often suffer from visual-linguistic biases present in datasets, which cause pre-trained vision-language models to overlook key details. To addres…