activity
20152023
most citedClothing Co-Parsing by Joint Image Segmentation and Labeling

154 citations · 848 across the 33 of their papers we have counts for

collaborators
Showing 2021Show all

12 papers · 1 filter

cs.CV202110 cited

Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and Language

Mingyu Ding, Zhenfang Chen, Tao Du +3

In this work, we propose a unified framework, called Visual Reasoning with Differ-entiable Physics (VRDP), that can jointly learn visual concepts and infer physics models of object…

cs.CV20216 cited

Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning

Chongjian Ge, Youwei Liang, Yibing Song +3

Studies on self-supervised visual representation learning (SSL) improve encoder backbones to discriminate training samples without labels. While CNN encoders via SSL achieve compar…

cs.CV20216 cited

Towards High-Quality Temporal Action Detection with Sparse Proposals

Jiannan Wu, Peize Sun, Shoufa Chen +4

Temporal Action Detection (TAD) is an essential and challenging topic in video understanding, aiming to localize the temporal segments containing human action instances and predict…

cs.LG20212 cited

Adversarial Robustness for Unsupervised Domain Adaptation

Muhammad Awais, Fengwei Zhou, Hang Xu +4

Extensive Unsupervised Domain Adaptation (UDA) studies have shown great success in practice by learning transferable representations across a labeled source domain and an unlabeled…

cs.CV20211 cited

Multi-frame Collaboration for Effective Endoscopic Video Polyp Detection via Spatial-Temporal Feature Transformation

Lingyun Wu, Zhiqiang Hu, Yuanfeng Ji +2

Precise localization of polyp is crucial for early cancer screening in gastrointestinal endoscopy. Videos given by endoscopy bring both richer contextual information as well as mor…

cs.CV202115 cited

Multi-Compound Transformer for Accurate Biomedical Image Segmentation

Yuanfeng Ji, Ruimao Zhang, Huijie Wang +4

The recent vision transformer(i.e.for image classification) learns non-local attentive interaction of different patch tokens. However, prior arts miss learning the cross-scale depe…