3 papers
cs.SD2026
DisSR: Disentangling Speech Representation for Degradation-Prior Guided Cross-Domain Speech Restoration
Ziqi Liang, Zhijun Jia, Chang Liu +3
Previous speech restoration (SR) primarily focuses on single-task speech restoration (SSR), which cannot address general speech restoration problems. Training specific SSR models f…
eess.AS2025
MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
Hongzhao Chen, XiaoYang Wang, Jing Lan +9
Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scar…
cs.CV2025
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
Tianqi Li, Ruobing Zheng, Minghui Yang +2
Recent advances in diffusion models have endowed talking head synthesis with subtle expressions and vivid head movements, but have also led to slow inference speed and insufficient…