87 citations · 152 across the 23 of their papers we have counts for
11 papers · 1 filter
Multi-Level Knowledge Distillation for Speech Emotion Recognition in Noisy Conditions
Yang Liu, Haoqin Sun, Geng Chen +4
Speech emotion recognition (SER) performance deteriorates significantly in the presence of noise, making it challenging to achieve competitive performance in noisy conditions. To t…
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
Cheng Gong, Xin Wang, Erica Cooper +5
Neural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis. Multilingual TTS systems are limited to resource-rich languages…
High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models
Chunyu Qiang, Hao Li, Yixin Tian +4
Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis deco…
Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding
Chunyu Qiang, Hao Li, Hao Ni +5
Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations a…
Rethinking the visual cues in audio-visual speaker extraction
Junjie Li, Meng Ge, Zexu pan +4
The Audio-Visual Speaker Extraction (AVSE) algorithm employs parallel video recording to leverage two visual cues, namely speaker identity and synchronization, to enhance performan…
speech and noise dual-stream spectrogram refine network with speech distortion loss for robust speech recognition
Haoyu Lu, Nan Li, Tongtong Song +4
In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. T…