activity
20162023
most citedMS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition

162 citations · 228 across the 11 of their papers we have counts for

collaborators

24 papers

eess.AS2023

OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition

Li Fu, Siqi Li, Qingtao Li +6

Self-Supervised Learning (SSL) Automatic Speech Recognition (ASR) models have shown great promise over Supervised Learning (SL) ones in low-resource settings. However, the advantag…

cs.CV2023★ 1 cited

Learning to Generate Poetic Chinese Landscape Painting with Calligraphy

Shaozu Yuan, Aijun Dai, Zhiling Yan +5

In this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generati…

cs.CV2022★ 27 cited

SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Huaishao Luo, Junwei Bao, Youzheng Wu +2

Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual…

cs.SD2022

MaskedSpeech: Context-aware Speech Synthesis with Masking Strategy

Ya-Jie Zhang, Wei Song, Yanghao Yue +3

Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis s…

cs.CL2022★ 2 cited

MoNET: Tackle State Momentum via Noise-Enhanced Training for Dialogue State Tracking

Haoning Zhang, Junwei Bao, Haipeng Sun +4

Dialogue state tracking (DST) aims to convert the dialogue history into dialogue states which consist of slot-value pairs. As condensed structural information memorizing all histor…

eess.AS2022

UFO2: A unified pre-training framework for online and offline speech recognition

Li Fu, Siqi Li, Qingtao Li +5

In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows…