Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Ziyang Ma, Zhikang Niu, Wenming Tu +30
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To sup…
cs.SD2026
SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
Hao Zhang, Yiwen Zhao, Yixuan Zhang +2
We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, a…
cs.SD2026
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…