collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS2025

Integrating IP Broadcasting with Audio Tags: Workflow and Challenges

Rhys Burchett-Vass, Arshdeep Singh, Gabriel Bibbó +1

The broadcasting industry has adopted IP technologies, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allo…

eess.AS2025

PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection

Jinbo Hu, Yin Cao, Ming Wu +5

Sound event localization and detection (SELD) has seen substantial advancements through learning-based methods. These systems, typically trained from scratch on specific datasets,…

eess.AS2025

Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities

Jinhua Liang, Xubo Liu, Wenwu Wang +3

The auditory system plays a substantial role in shaping the overall human perceptual experience. While prevailing large language models (LLMs) and visual language models (VLMs) hav…

eess.AS2024

Separate Anything You Describe

Xubo Liu, Qiuqiang Kong, Yan Zhao +7

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given…

eess.AS2024

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Jisheng Bai, Haohe Liu, Mou Wang +5

With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to th…

eess.AS2024

Universal Sound Separation with Self-Supervised Audio Masked Autoencoder

Junqi Zhao, Xubo Liu, Jinzheng Zhao +4

Universal sound separation (USS) is a task of separating mixtures of arbitrary sound sources. Typically, universal separation models are trained from scratch in a supervised manner…