collaborators

10 papers

cs.SD2025

Environmental Sound Classification on An Embedded Hardware Platform

Gabriel Bibbo, Arshdeep Singh, Mark D. Plumbley

Convolutional neural networks (CNNs) have exhibited state-of-the-art performance in various audio classification tasks. However, their real-time deployment remains a challenge on r…

eess.AS2025

Integrating IP Broadcasting with Audio Tags: Workflow and Challenges

Rhys Burchett-Vass, Arshdeep Singh, Gabriel Bibbó +1

The broadcasting industry has adopted IP technologies, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allo…

eess.AS2025

PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection

Jinbo Hu, Yin Cao, Ming Wu +5

Sound event localization and detection (SELD) has seen substantial advancements through learning-based methods. These systems, typically trained from scratch on specific datasets,…

eess.AS2025

Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities

Jinhua Liang, Xubo Liu, Wenwu Wang +3

The auditory system plays a substantial role in shaping the overall human perceptual experience. While prevailing large language models (LLMs) and visual language models (VLMs) hav…

cs.SD2025

FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

Yi Yuan, Xubo Liu, Haohe Liu +2

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches…

cs.SD2025

Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Yi Yuan, Dongya Jia, Xiaobin Zhuang +9

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performan…