48 citations · 277 across the 31 of their papers we have counts for
38 papers
Semi-supervised Multi-task Learning for Semantics and Depth
Yufeng Wang, Yi-Hsuan Tsai, Wei-Chih Hung +3
Multi-Task Learning (MTL) aims to enhance the model generalization by sharing representations between related tasks for better performance. Typical MTL methods are jointly trained…
Learning Cross-modal Contrastive Features for Video Domain Adaptation
Donghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang +4
Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation met…
Towards Interpretable Deep Networks for Monocular Depth Estimation
Zunzhi You, Yi-Hsuan Tsai, Wei-Chen Chiu +1
Deep networks for Monocular Depth Estimation (MDE) have achieved promising performance recently and it is of great importance to further understand the interpretability of these ne…
End-to-end Multi-modal Video Temporal Grounding
Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang
We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from…
Robust 360-8PA: Redesigning The Normalized 8-point Algorithm for 360-FoV Images
Bolivar Solarte, Chin-Hsuan Wu, Kuan-Wei Lu +3
This paper presents a novel preconditioning strategy for the classic 8-point algorithm (8-PA) for estimating an essential matrix from 360-FoV images (i.e., equirectangular images)…
Understanding Synonymous Referring Expressions via Contrastive Features
Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang
Referring expression comprehension aims to localize objects identified by natural language descriptions. This is a challenging task as it requires understanding of both visual and…