77 citations · 340 across the 50 of their papers we have counts for
3 papers · 1 filter
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
Andreea-Maria Oncescu, João F. Henriques, Andrew Zisserman +2
Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, trea…
Audio Retrieval with Natural Language Queries: A Benchmark Study
A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques +2
The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a gi…
Disentangled Speech Embeddings using Cross-modal Self-supervision
Arsha Nagrani, Joon Son Chung, Samuel Albanie +1
The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective tha…