3 papers
cs.SD2025
Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting
Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K
Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer…
cs.IR2025
Audio Prototypical Network For Controllable Music Recommendation
Fırat Öncel, Emiliano Penaloza, Haolun Wu +4
Traditional recommendation systems represent user preferences in dense representations obtained through black-box encoder models. While these models often provide strong recommenda…
cs.AI2025
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
Pooneh Mousavi, Shubham Gupta, Cem Subakan +1
Foundation models based on large language models (LLMs) have shown great success in handling various tasks and modalities. However, adapting these models for general-purpose audio-…