21 citations · 40 across the 5 of their papers we have counts for
19 papers
Efficient Cross-Modal Video Retrieval with Meta-Optimized Frames
Ning Han, Xun Yang, Ee-Peng Lim +2
Cross-modal video retrieval aims to retrieve the semantically relevant videos given a text as a query, and is one of the fundamental tasks in Multimedia. Most of top-performing met…
A Large-Scale Benchmark for Food Image Segmentation
Xiongwei Wu, Xin Fu, Ying Liu +3
Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmenta…
Counterfactual Zero-Shot and Open-Set Visual Recognition
Zhongqi Yue, Tan Wang, Hanwang Zhang +2
We present a novel counterfactual framework for both Zero-Shot Learning (ZSL) and Open-Set Recognition (OSR), whose common challenge is generalizing to the unseen-classes by only t…
Counterfactual Variable Control for Robust and Interpretable Question Answering
Sicheng Yu, Yulei Niu, Shuohang Wang +2
Deep neural network based question answering (QA) models are neither robust nor explainable in many cases. For example, a multiple-choice QA model, tested without any input of ques…
Causal Intervention for Weakly-Supervised Semantic Segmentation
Dong Zhang, Hanwang Zhang, Jinhui Tang +2
We present a causal inference framework to improve Weakly-Supervised Semantic Segmentation (WSSS). Specifically, we aim to generate better pixel-level pseudo-masks by using only im…
Interventional Few-Shot Learning
Zhongqi Yue, Hanwang Zhang, Qianru Sun +1
We uncover an ever-overlooked deficiency in the prevailing Few-Shot Learning (FSL) methods: the pre-trained knowledge is indeed a confounder that limits the performance. This findi…