4 citations · 10 across the 6 of their papers we have counts for
7 papers
Text-based Audio Retrieval by Learning from Similarities between Audio Captions
Huang Xie, Khazar Khorrami, Okko Räsänen +1
This paper proposes to use similarities of audio captions for estimating audio-caption relevances to be used for training text-based audio retrieval systems. Current audio-caption…
Multi-label Zero-Shot Audio Classification with Temporal Attention
Duygu Dogan, Huang Xie, Toni Heittola +1
Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot l…
Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
Huang Xie, Khazar Khorrami, Okko Räsänen +1
Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from…
Language-based Audio Retrieval Task in DCASE 2022 Challenge
Huang Xie, Samuel Lipping, Tuomas Virtanen
Language-based audio retrieval is a task, where natural language textual captions are used as queries to retrieve audio signals from a dataset. It has been first introduced into DC…
Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections
Huang Xie, Okko Räsänen, Tuomas Virtanen
In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Ze…
Zero-Shot Audio Classification via Semantic Embeddings
Huang Xie, Tuomas Virtanen
In this paper, we study zero-shot learning in audio classification via semantic embeddings extracted from textual labels and sentence descriptions of sound classes. Our goal is to…