10 citations · 23 across the 6 of their papers we have counts for
10 papers · 1 filter
AttrSeg: Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
Chaofan Ma, Yuhuan Yang, Chen Ju +3
Open-vocabulary semantic segmentation is a challenging task that requires segmenting novel object categories at inference time. Recent studies have explored vision-language pre-tra…
Multi-Modal Prototypes for Open-World Semantic Segmentation
Yuhuan Yang, Chaofan Ma, Chen Ju +4
In semantic segmentation, generalizing a visual system to both seen categories and novel categories at inference time has always been practically valuable yet challenging. To enabl…
Enhanced Multimodal Representation Learning with Cross-modal KD
Mengxi Chen, Linyu Xing, Yu Wang +1
This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Di…
Annotation-free Audio-Visual Segmentation
Jinxiang Liu, Yu Wang, Chen Ju +3
The objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task…
Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models
Chaofan Ma, Yuhuan Yang, Yanfeng Wang +2
When trained at a sufficient scale, self-supervised learning has exhibited a notable ability to solve a wide range of visual or language understanding tasks. In this paper, we inve…
Self-Supervised Masking for Unsupervised Anomaly Detection and Localization
Chaoqin Huang, Qinwei Xu, Yanfeng Wang +2
Recently, anomaly detection and localization in multimedia data have received significant attention among the machine learning community. In real-world applications such as medical…