3 papers
cs.SD2026
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
Ruinan Jin, Xinting Liao, Hanlin Yu +2
Modern voice cloning, also known as zero-shot text-to-speech (TTS), can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling appli…
eess.AS2025
Reference-aware SFM layers for intrusive intelligibility prediction
Hanlin Yu, Haoshuai Zhou, Boxuan Cao +3
Intrusive speech-intelligibility predictors that exploit explicit reference signals are now widespread, yet they have not consistently surpassed non-intrusive systems. We argue tha…
cs.SD2025
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
Boxuan Cao, Linkai Li, Hanlin Yu +3
Speech intelligibility evaluation for hearing-impaired (HI) listeners is essential for assessing hearing aid performance, traditionally relying on listening tests or intrusive meth…