30 citations · 46 across the 9 of their papers we have counts for
6 papers · 1 filter
Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
Kaifeng Gao, Long Chen, Hanwang Zhang +2
Prompt tuning with large-scale pretrained vision-language models empowers open-vocabulary predictions trained on limited base categories, e.g., object classification and detection.…
Equivariance and Invariance Inductive Bias for Learning from Insufficient Data
Tan Wang, Qianru Sun, Sugiri Pranata +2
We are interested in learning robust models from insufficient data, without the need for any externally pre-trained checkpoints. First, compared to sufficient data, we show why ins…
On Non-Random Missing Labels in Semi-Supervised Learning
Xinting Hu, Yulei Niu, Chunyan Miao +2
Semi-Supervised Learning (SSL) is fundamentally a missing label problem, in which the label Missing Not At Random (MNAR) problem is more realistic and challenging, compared to the…
RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Seas of videos are uploaded daily with the popularity of social channels; thus, retrieving the most related video contents with user textual queries plays a more crucial role. Most…
Deconfounded Visual Grounding
Jianqiang Huang, Yu Qin, Jiaxin Qi +2
We focus on the confounding bias between language and location in the visual grounding pipeline, where we find that the bias is the major visual reasoning bottleneck. For example,…
Cross-Domain Empirical Risk Minimization for Unbiased Long-tailed Classification
Beier Zhu, Yulei Niu, Xian-Sheng Hua +1
We address the overlooked unbiasedness in existing long-tailed classification methods: we find that their overall improvement is mostly attributed to the biased preference of tail…