activity
20162025
most citedHUME: Human UCCA-Based Evaluation of Machine Translation

8 citations · 19 across the 23 of their papers we have counts for

collaborators

23 papers

cs.CL2025

ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data

Vladislav Stankov, Matyáš Kopp, Ondřej Bojar

We introduce ParCzech4Speech 1.0, a processed version of the ParCzech 4.0 corpus, targeted at speech modeling tasks with the largest variant containing 2,695 hours. We combined the…

cs.CL2025

Preliminary Ranking of WMT25 General Machine Translation Systems

Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +25

We present the preliminary rankings of machine translation (MT) systems submitted to the WMT25 General Machine Translation Shared Task, as determined by automatic evaluation metric…

cs.CL2025

Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation

Petra Barančíková, Ondřej Bojar

In this paper, we compare Czech-specific and multilingual sentence embedding models through intrinsic and extrinsic evaluation paradigms. For intrinsic evaluation, we employ Costra…

cs.CL2025

Prompting LLMs: Length Control for Isometric Machine Translation

Dávid Javorský, Ondřej Bojar, François Yvon

In this study, we explore the effectiveness of isometric machine translation across multiple language pairs (EnDe, EnFr, and EnEs) under the conditions of the IWSLT…

cs.CL2025

MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines

Dávid Javorský, Ondřej Bojar, François Yvon

In simultaneous interpreting, an interpreter renders a source speech into another language with a very short lag, much sooner than sentences are finished. In order to understand an…

cs.CL2024

How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?

Sara Papi, Peter Polak, Ondřej Bojar +1

Simultaneous speech-to-text translation (SimulST) translates source-language speech into target-language text concurrently with the speaker's speech, ensuring low latency for bette…