10 papers · 1 filter
CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
Josef Jon, OndÅej Bojar
We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English…
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
Peter Polák, Sara Papi, Luisa Bentivogli +1
Simultaneous speech-to-text translation systems must balance translation quality with latency. Although quality evaluation is well established, latency measurement remains a challe…
Finetuning LLMs for EvaCun 2025 token prediction shared task
Josef Jon, OndÅej Bojar
In this paper, we present our submission for the token prediction task of EvaCun 2025. Our sys-tems are based on LLMs (Command-R, Mistral, and Aya Expanse) fine-tuned on the task d…
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
Nam Luu, OndÅej Bojar
Speech Translation (ST) is a machine translation task that involves converting speech signals from one language to the corresponding text in another language; this task has two dif…
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
Vladislav Stankov, Matyáš Kopp, OndÅej Bojar
We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the…
Preliminary Ranking of WMT25 General Machine Translation Systems
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +25
We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metric…