12 papers
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh +2
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real…
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
Pawat Chunhachatrachai, Gueter Josmy Faure, Hung-Ting Su +1
Spatial question answering over egocentric video is a challenging task that requires Vision-Language Models (VLMs) to reason about 3D object positions, scene affordances, and direc…
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
Posheng Chen, Powen Cheng, Gueter Josmy Faure +2
In real-world scenes, target objects may reside in regions that are not visible. While humans can often infer the locations of occluded objects from context and commonsense knowled…
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
Pei-An Chen, Yong-Ching Liang, Jia-Fong Yeh +4
Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. However, existing methods usually…
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
Hung-Ting Su, Ting-Jun Wang, Jia-Fong Yeh +2
Conventional Vision-and-Language Navigation (VLN) benchmarks assume instructions are feasible and the referenced target exists, leaving agents ill-equipped to handle false-premise…
Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
Tzu-Jung Lin, Jia-Fong Yeh, Hung-Ting Su +3
In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to…