6 papers
HoliAntiSpoof: Audio LLM for Holistic Speech Anti-Spoofing
Xuenan Xu, Yiming Ren, Liwei Liu +5
Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging. However, most existing methods treat spoofing as binary classification, overlook…
MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model
Ye Tao, Wen Wu, Chao Zhang +3
Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally li…
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
Qitan Lv, Tianyu Liu, Wen Wu +4
Speculative decoding (SD) has emerged as a promising approach to accelerate LLM inference without sacrificing output quality. Existing SD methods tailored for video-LLMs primarily…
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
Zihao Zheng, Zeyu Xie, Xuenan Xu +3
While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and in…
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
Xuenan Xu, Jiahao Mei, Zihao Zheng +9
Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where ea…
Can Audio Large Language Models Verify Speaker Identity?
Yiming Ren, Xuenan Xu, Baoxiang Li +2
This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive…