activity
20172021
most citedDeep Modular Co-Attention Networks for Visual Question Answering

99 citations · 160 across the 9 of their papers we have counts for

collaborators
Showing 2021Show all

6 papers · 1 filter

cs.CV2021

BiSTF: Bilateral-Branch Self-Training Framework for Semi-Supervised Large-scale Fine-Grained Recognition

Hao Chang, Guochen Xie, Jun Yu +1

Semi-supervised Fine-Grained Recognition is a challenge task due to the difficulty of data imbalance, high inter-class similarity and domain mismatch. Recent years, this field has…

cs.CV20211 cited

A Weakly-Supervised Depth Estimation Network Using Attention Mechanism

Fang Gao, Jiabao Wang, Jun Yu +2

Monocular depth estimation (MDE) is a fundamental task in many applications such as scene understanding and reconstruction. However, most of the existing methods rely on accurately…

cs.LG20217 cited

Instance Correction for Learning with Open-set Noisy Labels

Xiaobo Xia, Tongliang Liu, Bo Han +4

The problem of open-set noisy labels denotes that part of training data have a different label space that does not contain the true class. Lots of approaches, e.g., loss correction…

cs.LG202149 cited

Sample Selection with Uncertainty of Losses for Learning with Noisy Labels

Xiaobo Xia, Tongliang Liu, Bo Han +4

In learning with noisy labels, the sample selection approach is very popular, which regards small-loss data as correctly labeled during training. However, losses are generated on-t…

cs.CV20211 cited

Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching

Bofeng Wu, Guocheng Niu, Jun Yu +3

This paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved ta…

cs.LG2021

Removing Adversarial Noise in Class Activation Feature Space

Dawei Zhou, Nannan Wang, Chunlei Peng +4

Deep neural networks (DNNs) are vulnerable to adversarial noise. Preprocessing based defenses could largely remove adversarial noise by processing inputs. However, they are typical…