2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.SD2026
PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation
Lingfeng Yao, Chenpei Huang, Xingke Yang +5
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. How…
cs.SD2025★ 2 cited
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
Zihao Wang, Ruibin Yuan, Ziqi Geng +7
Automated singing assessment is crucial for education and entertainment. However, existing systems face two fundamental limitations: reliance on reference tracks, which stifles cre…