2 papers
cs.SD2026
Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a singl…
cs.CV2025
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
Ruixin Zhang, Jiaqing Fan, Yifan Liao +2
Referring Video Object Segmentation (RVOS) aims to segment specific objects in a video according to textual descriptions. We observe that recent RVOS approaches often place excessi…