28 citations · 38 across the 4 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023★ 3 cited
Efficient Audio Captioning Transformer with Patchout and Text Guidance
Thodoris Kouzelis, Grigoris Bastas, Athanasios Katsamanis +1
Automated audio captioning is multi-modal translation task that aim to generate textual descriptions for a given audio clip. In this paper we propose a full Transformer architectur…
cs.SD2023★ 28 cited
Designing and Evaluating Speech Emotion Recognition Systems: A reality check case study with IEMOCAP
Nikolaos Antoniou, Athanasios Katsamanis, Theodoros Giannakopoulos +1
There is an imminent need for guidelines and standard test sets to allow direct and fair comparisons of speech emotion recognition (SER). While resources, such as the Interactive E…