48 citations · 61 across the 18 of their papers we have counts for
18 papers
AudioSpa: Spatializing Sound Events with Text
Linfeng Feng, Lei Zhao, Boyu Zhu +2
Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text…
Why Does Dropping Edges Usually Outperform Adding Edges in Graph Contrastive Learning?
Yanchen Xu, Siqi Huang, Hongyuan Zhang +1
Graph contrastive learning (GCL) has been widely used as an effective self-supervised learning method for graph representation learning. However, how to apply adequate and stable g…
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation
Bingyu Li, Da Zhang, Zhiyuan Zhao +2
Open-vocabulary segmentation aims to identify and segment specific regions and objects based on text-based descriptions. A common solution is to leverage powerful vision-language m…
A Greedy Strategy for Graph Cut
Feiping Nie, Shenfei Pei, Zengwei Zheng +2
We propose a Greedy strategy to solve the problem of Graph Cut, called GGC. It starts from the state where each data sample is regarded as a cluster and dynamically merges the two…
Enhance Vision-Language Alignment with Noise
Sida Huang, Hongyuan Zhang, Xuelong Li
With the advancement of pre-trained vision-language (VL) models, enhancing the alignment between visual and linguistic modalities in downstream tasks has emerged as a critical chal…
SentenceVAE: Enable Next-sentence Prediction for Large Language Models with Faster Speed, Higher Accuracy and Longer Context
Hongjun An, Yifan Chen, Zhe Sun +1
Current large language models (LLMs) primarily utilize next-token prediction method for inference, which significantly impedes their processing speed. In this paper, we introduce a…