collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

Environmental Sound Classification on An Embedded Hardware Platform

Gabriel Bibbo, Arshdeep Singh, Mark D. Plumbley

Convolutional neural networks (CNNs) have exhibited state-of-the-art performance in various audio classification tasks. However, their real-time deployment remains a challenge on r…

cs.SD2025

FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

Yi Yuan, Xubo Liu, Haohe Liu +2

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches…

cs.SD2025

Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Yi Yuan, Dongya Jia, Xiaobin Zhuang +9

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performan…

cs.SD2024

SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Haohe Liu, Xuenan Xu, Yi Yuan +3

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelli…

cs.SD2024

The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection

Gabriel Bibbó, Thomas Deacon, Arshdeep Singh +1

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is…