1k citations · 1.8k across the 47 of their papers we have counts for
11 papers · 1 filter
Instance As Identity: A Generic Online Paradigm for Video Instance Segmentation
Feng Zhu, Zongxin Yang, Xin Yu +2
Modeling temporal information for both detection and tracking in a unified framework has been proved a promising solution to video instance segmentation (VIS). However, how to effe…
Label Semantic Knowledge Distillation for Unbiased Scene Graph Generation
Lin Li, Long Chen, Hanrong Shi +4
The Scene Graph Generation (SGG) task aims to detect all the objects and their pairwise visual relationships in a given image. Although SGG has achieved remarkable progress over th…
GPPF: A General Perception Pre-training Framework via Sparsely Activated Multi-Task Learning
Benyuan Sun, Jin Dai, Zihao Liang +3
Pre-training over mixtured multi-task, multi-domain, and multi-modal data remains an open challenge in vision perception pre-training. In this paper, we propose GPPF, a General Per…
Integrating Object-aware and Interaction-aware Knowledge for Weakly Supervised Scene Graph Generation
Xingchen Li, Long Chen, Wenbo Ma +2
Recently, increasing efforts have been focused on Weakly Supervised Scene Graph Generation (WSSGG). The mainstream solution for WSSGG typically follows the same pipeline: they firs…
Subband-based Generative Adversarial Network for Non-parallel Many-to-many Voice Conversion
Jian Ma, Zhedong Zheng, Hao Fei +3
Voice conversion is to generate a new speech with the source content and a target voice style. In this paper, we focus on one general setting, i.e., non-parallel many-to-many voice…
VL: Leveraging Vision and Vision-language Models into Large-scale Product Retrieval
Wenhao Wang, Yifan Sun, Zongxin Yang +1
Product retrieval is of great importance in the ecommerce domain. This paper introduces our 1st-place solution in eBay eProduct Visual Search Challenge (FGVC9), which is featured f…