Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
Scaling Audio-Text Retrieval with Multimodal Large Language Models
Jilan Xu, Carl Thomé, Danijela Horak +2
Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentall…
cs.SD2023
Retrieval Augmented Generation of Symbolic Music with LLMs
Nicolas Jonason, Luca Casini, Carl Thomé +1
We explore the use of large language models (LLMs) for music generation using a retrieval system to select relevant examples. We find promising initial results for music generation…
cs.SD2020
Perceiving Music Quality with GANs
Agrin Hilmkil, Carl Thomé, Anders Arpteg
Several methods have been developed to assess the perceptual quality of audio under transforms like lossy compression. However, they require paired reference signals of the unalter…