activity
20192024
most citedZero-Shot Audio Classification via Semantic Embeddings

4 citations · 10 across the 6 of their papers we have counts for

collaborators

7 papers

eess.AS2024

Text-based Audio Retrieval by Learning from Similarities between Audio Captions

Huang Xie, Khazar Khorrami, Okko Räsänen +1

This paper proposes to use similarities of audio captions for estimating audio-caption relevances to be used for training text-based audio retrieval systems. Current audio-caption…

cs.SD2024

Multi-label Zero-Shot Audio Classification with Temporal Attention

Duygu Dogan, Huang Xie, Toni Heittola +1

Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot l…

eess.AS2024

Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning

Huang Xie, Khazar Khorrami, Okko Räsänen +1

Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from…

eess.AS20224 cited

Language-based Audio Retrieval Task in DCASE 2022 Challenge

Huang Xie, Samuel Lipping, Tuomas Virtanen

Language-based audio retrieval is a task, where natural language textual captions are used as queries to retrieve audio signals from a dataset. It has been first introduced into DC…

eess.AS20202 cited

Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections

Huang Xie, Okko Räsänen, Tuomas Virtanen

In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Ze…

eess.AS20204 cited

Zero-Shot Audio Classification via Semantic Embeddings

Huang Xie, Tuomas Virtanen

In this paper, we study zero-shot learning in audio classification via semantic embeddings extracted from textual labels and sentence descriptions of sound classes. Our goal is to…