10 citations · 11 across the 3 of their papers we have counts for
4 papers
Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling
Hsin-Ying Lee, Hung-Ting Su, Bing-Chen Tsai +3
While recent large-scale video-language pre-training made great progress in video question answering, the design of spatial modeling of video-language models is less fine-grained t…
MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer
Kuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su +1
Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to…
: Learnable Sparse Signal Superdensity for Guided Depth Estimation
Yu-Kai Huang, Yueh-Cheng Liu, Tsung-Han Wu +5
Dense depth estimation plays a key role in multiple applications such as robotics, 3D reconstruction, and augmented reality. While sparse signal, e.g., LiDAR and Radar, has been le…
Expanding Sparse Guidance for Stereo Matching
Yu-Kai Huang, Yueh-Cheng Liu, Tsung-Han Wu +2
The performance of image based stereo estimation suffers from lighting variations, repetitive patterns and homogeneous appearance. Moreover, to achieve good performance, stereo sup…