activity
20242026
collaborators

11 papers

cs.CL2026

CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

Josef Jon, Ondřej Bojar

We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English…

cs.HC2026

MEEDAV: A Synchronous Web Viewer for EEG, Eye-Tracking and Speech Data

Jan Pijálek, Karel Vlk, Ondřej Bojar

MEEDAV is an open-source web-based application for the synchronised visualisation of electroencephalography (EEG), eye-tracking, and audio data collected in psycholinguistic resear…

cs.CL2026

Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation

Peter Polák, Sara Papi, Luisa Bentivogli +1

Simultaneous speech-to-text translation systems must balance translation quality with latency. Although quality evaluation is well established, latency measurement remains a challe…

cs.CL2025

Finetuning LLMs for EvaCun 2025 token prediction shared task

Josef Jon, Ondřej Bojar

In this paper, we present our submission for the token prediction task of EvaCun 2025. Our sys-tems are based on LLMs (Command-R, Mistral, and Aya Expanse) fine-tuned on the task d…

cs.CL2025

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

Nam Luu, Ondřej Bojar

Speech Translation (ST) is a machine translation task that involves converting speech signals from one language to the corresponding text in another language; this task has two dif…

cs.CL2025

ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data

Vladislav Stankov, Matyáš Kopp, Ondřej Bojar

We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the…