2 citations · 3 across the 5 of their papers we have counts for
5 papers
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
Tan Dat Nguyen, Ji-Hoon Kim, Youngjoon Jang +2
The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key componen…
Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for Tempo Prediction and Search
Matthew C. McCallum, Florian Henkel, Jaehun Kim +2
Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it…
On the Effect of Data-Augmentation on Local Embedding Properties in the Contrastive Learning of Music Audio Representations
Matthew C. McCallum, Matthew E. P. Davies, Florian Henkel +2
Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of down…
Seeing Through the Conversation: Audio-Visual Speech Separation based on Diffusion Model
Suyeon Lee, Chaeyoung Jung, Youngjoon Jang +2
The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their…
Contrastive Learning for Cross-modal Artist Retrieval
Andres Ferraro, Jaehun Kim, Sergio Oramas +2
Music retrieval and recommendation applications often rely on content features encoded as embeddings, which provide vector representations of items in a music dataset. Numerous com…