23 citations · 127 across the 27 of their papers we have counts for
3 papers · 1 filter
Language Models are Causal Knowledge Extractors for Zero-shot Video Question Answering
Hung-Ting Su, Yulei Niu, Xudong Lin +2
Causal Video Question Answering (CVidQA) queries not only association or temporal relations but also causal relations in a video. Existing question synthesis methods pre-trained qu…
OCID-Ref: A 3D Robotic Dataset with Embodied Language for Clutter Scene Grounding
Ke-Jyun Wang, Yun-Hsuan Liu, Hung-Ting Su +4
To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded…
Situation and Behavior Understanding by Trope Detection on Films
Chen-Hsi Chang, Hung-Ting Su, Jui-heng Hsu +7
The human ability of deep cognitive skills are crucial for the development of various real-world applications that process diverse and abundant user generated input. While recent p…