5 papers
Multilingual Audio Captioning using machine translated data
Matéo Cousin, Étienne Labbé, Thomas Pellegrini
Automated Audio Captioning (AAC) systems attempt to generate a natural language sentence, a caption, that describes the content of an audio recording, in terms of sound events. Exi…
CoNeTTE: An efficient Audio Captioning system leveraging multiple datasets with Task Embedding
Étienne Labbé, Thomas Pellegrini, Julien Pinquier
Automated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures. An audio encoder produces audio embeddings…
Killing two birds with one stone: Can an audio captioning system also be used for audio-text retrieval?
Etienne Labbé, Thomas Pellegrini, Julien Pinquier
Automated Audio Captioning (AAC) aims to develop systems capable of describing an audio recording using a textual sentence. In contrast, Audio-Text Retrieval (ATR) systems seek to…
Adapting a ConvNeXt model to audio classification on AudioSet
Thomas Pellegrini, Ismail Khalfaoui-Hassani, Etienne Labbé +1
In computer vision, convolutional neural networks (CNN) such as ConvNeXt, have been able to surpass state-of-the-art transformers, partly thanks to depthwise separable convolutions…
Multitask learning in Audio Captioning: a sentence embedding regression loss acts as a regularizer
Etienne Labbé, Julien Pinquier, Thomas Pellegrini
In this work, we propose to study the performance of a model trained with a sentence embedding regression loss component for the Automated Audio Captioning task. This task aims to…