most citedSPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2026

MMAG: A Multi-Control Mixed Audio Generation Benchmark

Zihao Zheng, Xuenan Xu, Jiahao Mei +5

Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluat…

cs.SD2026

Bayesian Speech Synthesizers Can Learn from Multiple Teachers

Ziyang Zhang, Yifan Gao, Xuenan Xu +3

Text-to-Speech (TTS) is inherently a "one-to-many" mapping characterized by intrinsic uncertainty, yet current paradigms often oversimplify it into a deterministic regression task.…

cs.SD2026

HoliAntiSpoof: Audio LLM for Holistic Speech Anti-Spoofing

Xuenan Xu, Yiming Ren, Liwei Liu +5

Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging. However, most existing methods treat spoofing as binary classification, overlook…

cs.SD2026

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

Ye Tao, Wen Wu, Chao Zhang +3

Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally li…

cs.SD2025

PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description

Zihao Zheng, Zeyu Xie, Xuenan Xu +3

While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and in…

cs.SD2025

UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities

Xuenan Xu, Jiahao Mei, Zihao Zheng +9

Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where ea…