2 papers
cs.CL2026
Learning When to Think While Listening in Large Audio-Language Models
Zhiyuan Song, Weici Zhao, Yang Xiao +3
Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsivene…
cs.SD2025
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
Fu Li, Weichao Zhao, You Li +2
Recent advances in video generation have achieved remarkable improvements in visual content fidelity. However, the absence of synchronized audio severely undermines immersive exper…