activity
20242026
collaborators

14 papers

cs.SD2026

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

Yi Yuan, Xubo Liu, Haohe Liu +5

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite produc…

eess.AS2026

SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training

Xinhao Mei, Gael Le Lan, Haohe Liu +5

Contrastive language-audio pretraining (CLAP) has achieved notable success in learning semantically rich audio representations and is widely adopted for various audio-related tasks…

cs.SD2025

EnvSDD: Benchmarking Environmental Sound Deepfake Detection

Han Yin, Yang Xiao, Rohan Kumar Das +4

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech o…

eess.AS2025

Region-Specific Audio Tagging for Spatial Sound

Jinzheng Zhao, Yong Xu, Haohe Liu +6

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given r…

cs.SD2025

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

Junqi Zhao, Jinzheng Zhao, Haohe Liu +5

Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by lear…

eess.AS2025

Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows

Haohe Liu, Thomas Deacon, Wenwu Wang +2

Locating the right sound effect efficiently is an important yet challenging topic for audio production. Most current sound-searching systems rely on pre-annotated audio labels crea…