23 citations · 51 across the 9 of their papers we have counts for
12 papers
Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling
Hsin-Ying Lee, Hung-Ting Su, Bing-Chen Tsai +3
While recent large-scale video-language pre-training made great progress in video question answering, the design of spatial modeling of video-language models is less fine-grained t…
MonoDTR: Monocular 3D Object Detection with Depth-Aware Transformer
Kuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su +1
Monocular 3D object detection is an important yet challenging task in autonomous driving. Some existing methods leverage depth information from an off-the-shelf depth estimator to…
Multivariate and Propagation Graph Attention Network for Spatial-Temporal Prediction with Outdoor Cellular Traffic
Chung-Yi Lin, Hung-Ting Su, Shen-Lung Tung +1
Spatial-temporal prediction is a critical problem for intelligent transportation, which is helpful for tasks such as traffic control and accident prevention. Previous studies rely…
TrUMAn: Trope Understanding in Movies and Animations
Hung-Ting Su, Po-Wei Shen, Bing-Chen Tsai +3
Understanding and comprehending video content is crucial for many real-world applications such as search and recommendation systems. While recent progress of deep learning has boos…
OCID-Ref: A 3D Robotic Dataset with Embodied Language for Clutter Scene Grounding
Ke-Jyun Wang, Yun-Hsuan Liu, Hung-Ting Su +4
To effectively apply robots in working environments and assist humans, it is essential to develop and evaluate how visual grounding (VG) can affect machine performance on occluded…
: Learnable Sparse Signal Superdensity for Guided Depth Estimation
Yu-Kai Huang, Yueh-Cheng Liu, Tsung-Han Wu +5
Dense depth estimation plays a key role in multiple applications such as robotics, 3D reconstruction, and augmented reality. While sparse signal, e.g., LiDAR and Radar, has been le…