activity
20172021
most citedDeep Modular Co-Attention Networks for Visual Question Answering

99 citations · 160 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2021

BiSTF: Bilateral-Branch Self-Training Framework for Semi-Supervised Large-scale Fine-Grained Recognition

Hao Chang, Guochen Xie, Jun Yu +1

Semi-supervised Fine-Grained Recognition is a challenge task due to the difficulty of data imbalance, high inter-class similarity and domain mismatch. Recent years, this field has…

cs.CV20211 cited

A Weakly-Supervised Depth Estimation Network Using Attention Mechanism

Fang Gao, Jiabao Wang, Jun Yu +2

Monocular depth estimation (MDE) is a fundamental task in many applications such as scene understanding and reconstruction. However, most of the existing methods rely on accurately…

cs.CV20211 cited

Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching

Bofeng Wu, Guocheng Niu, Jun Yu +3

This paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved ta…

cs.CV20203 cited

Learning Domain-invariant Graph for Adaptive Semi-supervised Domain Adaptation with Few Labeled Source Samples

Jinfeng Li, Weifeng Liu, Yicong Zhou +2

Domain adaptation aims to generalize a model from a source domain to tackle tasks in a related but different target domain. Traditional domain adaptation algorithms assume that eno…

cs.CV2020

Weakly-Supervised Multi-Level Attentional Reconstruction Network for Grounding Textual Queries in Videos

Yijun Song, Jingwen Wang, Lin Ma +2

The task of temporally grounding textual queries in videos is to localize one video segment that semantically corresponds to the given query. Most of the existing approaches rely o…

cs.CV2019

Multimodal Unified Attention Networks for Vision-and-Language Interactions

Zhou Yu, Yuhao Cui, Jun Yu +2

Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual co…