9 citations · 11 across the 3 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2026
Vclip: Face-based Speaker Generation by Face-voice Association Learning
Yao Shi, Yunfei Xu, Hongbin Suo +2
This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a refe…
eess.AS2025
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
Cancan Li, Fei Su, Juan Liu +4
Whisper speech recognition is crucial not only for ensuring privacy in sensitive communications but also for providing a critical communication bridge for patients under vocal rest…
eess.AS2023
Task-Agnostic Structured Pruning of Speech Representation Models
Haoyu Wang, Siyuan Wang, Wei-Qiang Zhang +2
Self-supervised pre-trained models such as Wav2vec2, Hubert, and WavLM have been shown to significantly improve many speech tasks. However, their large memory and strong computatio…