23 citations · 57 across the 7 of their papers we have counts for
7 papers
End-to-End Video Question-Answer Generation with Generator-Pretester Network
Hung-Ting Su, Chen-Hsi Chang, Po-Wei Shen +5
We study a novel task, Video Question-Answer Generation (VQAG), for challenging Video Question Answering (Video QA) task in multimedia. Due to expensive data annotation costs, many…
Situation and Behavior Understanding by Trope Detection on Films
Chen-Hsi Chang, Hung-Ting Su, Jui-heng Hsu +7
The human ability of deep cognitive skills are crucial for the development of various real-world applications that process diverse and abundant user generated input. While recent p…
GDN: A Coarse-To-Fine (C2F) Representation for End-To-End 6-DoF Grasp Detection
Kuang-Yu Jeng, Yueh-Cheng Liu, Zhe Yu Liu +4
We proposed an end-to-end grasp detection network, Grasp Detection Network (GDN), cooperated with a novel coarse-to-fine (C2F) grasp representation design to detect diverse and acc…
Unified Representation Learning for Cross Model Compatibility
Chien-Yi Wang, Ya-Liang Chang, Shang-Ta Yang +2
We propose a unified representation learning framework to address the Cross Model Compatibility (CMC) problem in the context of visual search applications. Cross compatibility betw…
Deep Long Audio Inpainting
Ya-Liang Chang, Kuan-Ying Lee, Po-Yu Wu +2
Long (> 200 ms) audio inpainting, to recover a long missing part in an audio segment, could be widely applied to audio editing tasks and transmission loss recovery. It is a very ch…
VORNet: Spatio-temporally Consistent Video Inpainting for Object Removal
Ya-Liang Chang, Zhe Yu Liu, Winston Hsu
Video object removal is a challenging task in video processing that often requires massive human efforts. Given the mask of the foreground object in each frame, the goal is to comp…