162 citations · 228 across the 11 of their papers we have counts for
24 papers
OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition
Li Fu, Siqi Li, Qingtao Li +6
Self-Supervised Learning (SSL) Automatic Speech Recognition (ASR) models have shown great promise over Supervised Learning (SL) ones in low-resource settings. However, the advantag…
Learning to Generate Poetic Chinese Landscape Painting with Calligraphy
Shaozu Yuan, Aijun Dai, Zhiling Yan +5
In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generati…
SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation
Huaishao Luo, Junwei Bao, Youzheng Wu +2
Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual…
MaskedSpeech: Context-aware Speech Synthesis with Masking Strategy
Ya-Jie Zhang, Wei Song, Yanghao Yue +3
Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis s…
MoNET: Tackle State Momentum via Noise-Enhanced Training for Dialogue State Tracking
Haoning Zhang, Junwei Bao, Haipeng Sun +4
Dialogue state tracking (DST) aims to convert the dialogue history into dialogue states which consist of slot-value pairs. As condensed structural information memorizing all histor…
UFO2: A unified pre-training framework for online and offline speech recognition
Li Fu, Siqi Li, Qingtao Li +5
In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows…