17 citations · 20 across the 7 of their papers we have counts for
7 papers
Advancing Referring Expression Segmentation Beyond Single Image
Yixuan Wu, Zhao Zhang, Xie Chi +2
Referring Expression Segmentation (RES) is a widely explored multi-modal task, which endeavors to segment the pre-existing object within a single image with a given linguistic expr…
CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-Matching
Xiaoshi Wu, Feng Zhu, Rui Zhao +1
Open-vocabulary detection (OVD) is an object detection task aiming at detecting objects from novel categories beyond the base categories on which the detector is trained. Recent OV…
HumanBench: Towards General Human-centric Perception with Projector Assisted Pretraining
Shixiang Tang, Cheng Chen, Qingsong Xie +9
Human-centric perceptions include a variety of vision tasks, which have widespread industrial applications, including surveillance, autonomous driving, and the metaverse. It is des…
Saliency Guided Contrastive Learning on Scene Images
Meilin Chen, Yizhou Wang, Shixiang Tang +6
Self-supervised learning holds promise in leveraging large numbers of unlabeled data. However, its success heavily relies on the highly-curated dataset, e.g., ImageNet, which still…
DCMT: A Direct Entire-Space Causal Multi-Task Framework for Post-Click Conversion Estimation
Feng Zhu, Mingjie Zhong, Xinxing Yang +8
In recommendation scenarios, there are two long-standing challenges, i.e., selection bias and data sparsity, which lead to a significant drop in prediction accuracy for both Click-…
Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation
Feng Zhu, Zongxin Yang, Xin Yu +2
Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effe…