2 citations · 3 across the 9 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
MMAG: A Multi-Control Mixed Audio Generation Benchmark
Zihao Zheng, Xuenan Xu, Jiahao Mei +5
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…
cs.SD2026
OCR-Enhanced Multimodal ASR Can Read While Listening
Junli Chen, Changli Tang, Yixuan Li +2
Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to…