4 papers
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski +4
The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identific…
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
Marek Kubis, Paweł Skórzewski, Iwona Christop +4
The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes tex…
CAMEO: Collection of Multilingual Emotional Speech Corpora
Iwona Christop, Maciej Czajka
This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. T…
ClonEval: An Open Voice Cloning Benchmark
Iwona Christop, Tomasz Kuczyński, Marek Kubis
We present a novel benchmark for voice cloning text-to-speech models. The benchmark consists of an evaluation protocol, an open-source library for assessing the performance of voic…