3 papers
eess.AS2024
Text-based Audio Retrieval by Learning from Similarities between Audio Captions
Huang Xie, Khazar Khorrami, Okko Räsänen +1
This paper proposes to use similarities of audio captions for estimating audio-caption relevances to be used for training text-based audio retrieval systems. Current audio-caption…
eess.AS2024
Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning
Huang Xie, Khazar Khorrami, Okko Räsänen +1
Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from…
eess.AS2024
A model of early word acquisition based on realistic-scale audiovisual naming events
Khazar Khorrami, Okko Räsänen
Infants gradually learn to parse continuous speech into words and connect names with objects, yet the mechanisms behind development of early word perception skills remain unknown.…