2 papers
eess.AS2026
FSD50K-Solo: Automated Curation of Single-Source Sound Events
Ningyuan Yang, Sile Yin, Li-Chia Yang +4
High-quality training datasets are essential for the performance of neural networks. However, the audio domain still lacks a large-scale, strongly-labeled, and single-source sound…
cs.SD2026
No Word Left Behind: Mitigating Prefix Bias in Open-Vocabulary Keyword Spotting
Yi Liu, Chuan-Che Huang, Xiao Quan
Open-vocabulary keyword spotting (OV-KWS) enables personalized device control via arbitrary voice commands. Recently, researchers have explored using audio-text joint embeddings, a…