most citedContrastive Learning for Cross-modal Artist Retrieval

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder

Tan Dat Nguyen, Ji-Hoon Kim, Youngjoon Jang +2

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key componen…

cs.SD2024

Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search

Matthew C. McCallum, Florian Henkel, Jaehun Kim +2

Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it…

cs.SD2024

On the Effect of Data-Augmentation on Local Embedding Properties in the Contrastive Learning of Music Audio Representations

Matthew C. McCallum, Matthew E. P. Davies, Florian Henkel +2

Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of down…

eess.AS20231 cited

Seeing Through the Conversation: Audio-Visual Speech Separation based on Diffusion Model

Suyeon Lee, Chaeyoung Jung, Youngjoon Jang +2

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their…

cs.IR20232 cited

Contrastive Learning for Cross-modal Artist Retrieval

Andres Ferraro, Jaehun Kim, Sergio Oramas +2

Music retrieval and recommendation applications often rely on content features encoded as embeddings, which provide vector representations of items in a music dataset. Numerous com…