117 citations · 233 across the 42 of their papers we have counts for
38 papers · 1 filter
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Kai Li, Wenze Ren, Junjie Li +11
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols c…
Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN
Yin-Ping Cho, Yu Tsao, Hsin-Min Wang +1
Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this wor…
Partial Coupling of Optimal Transport for Spoken Language Identification
Xugang Lu, Peng Shen, Yu Tsao +1
In order to reduce domain discrepancy to improve the performance of cross-domain spoken language identification (SLID) system, as an unsupervised domain adaptation (UDA) method, we…
Continuous Speech for Improved Learning Pathological Voice Disorders
Syu-Siang Wang, Chi-Te Wang, Chih-Chung Lai +2
Goal: Numerous studies had successfully differentiated normal and abnormal voice samples. Nevertheless, further classification had rarely been attempted. This study proposes a nove…
A Speech Intelligibility Enhancement Model based on Canonical Correlation and Deep Learning for Hearing-Assistive Technologies
Tassadaq Hussain, Muhammad Diyan, Mandar Gogate +4
Current deep learning (DL) based approaches to speech intelligibility enhancement in noisy environments are generally trained to minimise the distance between clean and enhanced sp…
EMGSE: Acoustic/EMG Fusion for Multimodal Speech Enhancement
Kuan-Chen Wang, Kai-Chun Liu, Hsin-Min Wang +1
Multimodal learning has been proven to be an effective method to improve speech enhancement (SE) performance, especially in challenging situations such as low signal-to-noise ratio…