2 citations · 3 across the 3 of their papers we have counts for
5 papers
OCID-Ref: A 3D Robotic Dataset with Embodied Language for Clutter Scene Grounding
Ke-Jyun Wang, Yun-Hsuan Liu, Hung-Ting Su +4
To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded…
End-to-End Video Question-Answer Generation with Generator-Pretester Network
Hung-Ting Su, Chen-Hsi Chang, Po-Wei Shen +5
We study a novel task, Video Question-Answer Generation (VQAG), for challenging Video Question Answering (Video QA) task in multimedia. Due to expensive data annotation costs, many…
Situation and Behavior Understanding by Trope Detection on Films
Chen-Hsi Chang, Hung-Ting Su, Jui-heng Hsu +7
The human ability of deep cognitive skills are crucial for the development of various real-world applications that process diverse and abundant user generated input. While recent p…
Investigating the Decoders of Maximum Likelihood Sequence Models: A Look-ahead Approach
Yu-Siang Wang, Yen-Ling Kuo, Boris Katz
We demonstrate how we can practically incorporate multi-step future information into a decoder of maximum likelihood sequence models. We propose a "k-step look-ahead" module to con…
Video Question Generation via Cross-Modal Self-Attention Networks Learning
Yu-Siang Wang, Hung-Ting Su, Chen-Hsi Chang +2
We introduce a novel task, Video Question Generation (Video QG). A Video QG model automatically generates questions given a video clip and its corresponding dialogues. Video QG req…