3 papers
cs.SD2026
NVV-SuperBench: Beyond Words, Beyond Quality-Benchmarking Nonverbal Vocalizations in Speech Generation
Liumeng Xue, Weizhen Bian, Jiahao Pan +9
Nonverbal vocalizations (NVVs), such as laughing, sighing, and sobbing, are essential for human-like speech, yet standardized evaluation rarely jointly assesses whether systems gen…
cs.CL2025
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
Anna Min, Chenxu Hu, Yi Ren +1
Though end-to-end speech-to-text translation has been a great success, we argue that the cascaded speech-to-text translation model still has its place, which is usually criticized…
cs.CL2025
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
Anna Min, Chenxu Hu, Yi Ren +1
Current research in speech-to-speech translation (S2ST) primarily concentrates on translation accuracy and speech naturalness, often overlooking key elements like paralinguistic in…