23 citations · 51 across the 26 of their papers we have counts for
13 papers · 1 filter
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
Gueter Josmy Faure, Min-Hung Chen, Jia-Fong Yeh +2
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real…
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
Pawat Chunhachatrachai, Gueter Josmy Faure, Hung-Ting Su +1
Spatial question answering over egocentric video is a challenging task that requires Vision-Language Models (VLMs) to reason about 3D object positions, scene affordances, and direc…
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
Posheng Chen, Powen Cheng, Gueter Josmy Faure +2
In real-world scenes, target objects may reside in regions that are not visible. While humans can often infer the locations of occluded objects from context and commonsense knowled…
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
Gueter Josmy Faure, Jia-Fong Yeh, Min-Hung Chen +3
Long-form video understanding presents unique challenges that extend beyond traditional short-video analysis approaches, particularly in capturing long-range dependencies, processi…
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
Hung-Ting Su, Chun-Tong Chao, Ya-Ching Hsu +4
Large Language Models (LLMs) have demonstrated effectiveness not only in language tasks but also in video reasoning. This paper introduces a novel dataset, Tropes in Movies (TiM),…
Tel2Veh: Fusion of Telecom Data and Vehicle Flow to Predict Camera-Free Traffic via a Spatio-Temporal Framework
ChungYi Lin, Shen-Lung Tung, Hung-Ting Su +1
Vehicle flow, a crucial indicator for transportation, is often limited by detector coverage. With the advent of extensive mobile network coverage, we can leverage mobile user activ…