most citedRep2wav: Noise Robust text-to-speech Using self-supervised representations

1 citations · 3 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV20241 cited

Adaptive Confidence Multi-View Hashing for Multimedia Retrieval

Jian Zhu, Yu Cui, Zhangmin Huang +4

The multi-view hash method converts heterogeneous data from multiple views into binary hash codes, which is one of the critical technologies in multimedia retrieval. However, the c…

eess.AS2024

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

Qiushi Zhu, Jie Zhang, Yu Gu +2

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-fi…

eess.AS20231 cited

Rep2wav: Noise Robust text-to-speech Using self-supervised representations

Qiushi Zhu, Yu Gu, Rilin Chen +4

Benefiting from the development of deep learning, text-to-speech (TTS) techniques using clean speech have achieved significant performance improvements. The data collected from rea…

eess.AS2023

CASA-ASR: Context-Aware Speaker-Attributed ASR

Mohan Shi, Zhihao Du, Qian Chen +5

Recently, speaker-attributed automatic speech recognition (SA-ASR) has attracted a wide attention, which aims at answering the question ``who spoke what''. Different from modular s…

eess.AS20231 cited

Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction

Mohan Shi, Yuchun Shu, Lingyun Zuo +4

For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach…

eess.AS2023

Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection

Xiao-Min Zeng, Yan Song, Zhu Zhuo +5

In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE)…