2 papers
cs.SD2026
EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
Xin Jing, Andreas Triantafyllopoulos, Jiadong Wang +3
Recent advancements in speech captioning models have enabled the generation of rich, fine-grained captions for emotional speech. However, the evaluation of such captions remains a…
eess.AS2025
MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
Hongzhao Chen, XiaoYang Wang, Jing Lan +9
Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scar…