11 papers
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
Josef Jon, OndÅej Bojar
We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English…
MEEDAV: A Synchronous Web Viewer for EEG, Eye-Tracking and Speech Data
Jan Pijálek, Karel Vlk, OndÅej Bojar
MEEDAV is an open-source web-based application for the synchronised visualisation of electroencephalography (EEG), eye-tracking, and audio data collected in psycholinguistic resear…
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
Peter Polák, Sara Papi, Luisa Bentivogli +1
Simultaneous speech-to-text translation systems must balance translation quality with latency. Although quality evaluation is well established, latency measurement remains a challe…
Finetuning LLMs for EvaCun 2025 token prediction shared task
Josef Jon, OndÅej Bojar
In this paper, we present our submission for the token prediction task of EvaCun 2025. Our sys-tems are based on LLMs (Command-R, Mistral, and Aya Expanse) fine-tuned on the task d…
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
Nam Luu, OndÅej Bojar
Speech Translation (ST) is a machine translation task that involves converting speech signals from one language to the corresponding text in another language; this task has two dif…
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Vladislav Stankov, Matyáš Kopp, OndÅej Bojar
We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the…