works on

From the 2 of 17 linked papers with an AI index.

collaborators

17 papers

cs.SD2026

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

Wenxiang Guo, Changhao Pan, Ziyue Jiang +2

Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast producti…

eess.AS2026

Qwen-Audio-VAE Technical Report

Ziyue Jiang, Dake Guo, Zekai Zhang +11

Qwen-Audio-VAE is a low‑bitrate, fast‑encoding continuous audio autoencoder that produces compact latent representations for scalable text‑to‑audio generation, using a causal encod…

cs.SD2026

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan +24

Qwen-Music is a large language model‑based system that generates high‑fidelity songs with vocals from text prompts or re‑imagines existing tracks, using a semantic token representa…

eess.AS2026

Audio Editing in the Era of Foundation Models: A Survey

Changhao Pan, Yifei Fan, Fan Zhuo +13

Audio editing aims to modify a given synthetic or real-world audio signal to satisfy specific user needs. As a promising yet challenging direction in AIGC, it has attracted increas…

eess.AS2026

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13

Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial rel…

eess.AS2026

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

Ke Lei, Yu Zhang, Changhao Pan +4

Real-time and accurate spatial audio generation is pivotal for delivering an immersive experience. However, existing spatial audio synthesis technologies are often encumbered by a…