Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski +4
The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identific…
cs.SD2025
LoRP-TTS: Low-Rank Personalized Text-To-Speech
Łukasz Bondaruk, Jakub Kubiak
Speech synthesis models convert written text into natural-sounding audio. While earlier models were limited to a single speaker, recent advancements have led to the development of…