6 papers
Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages
Sofiia Riazhskykh, Nam Luu, Ondřej Bojar
Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed t…
Dynamically Allocating Evaluation Effort for Model Ranking
Vilém Zouhar, Vilém Zouhar, Julia Kreutzer +6
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation pr…
Findings of the IWSLT 2024 Evaluation Campaign
Ibrahim Said Ahmad, Antonios Anastasopoulos, OndÅej Bojar +42
This paper reports on the shared tasks organized by the 21st IWSLT Conference. The shared tasks address 7 scientific challenges in spoken language translation: simultaneous and off…
Adversarial Testing as a Tool for Interpretability: Length-based Overfitting of Elementary Functions in Transformers
Patrik Zavoral, DuÅ¡an VariÅ¡, OndÅej Bojar
The Transformer model has a tendency to overfit various aspects of the training data, such as the overall sequence length. We study elementary string edit functions using a defined…
Preliminary WMT24 Ranking of General MT Systems and LLMs
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden +18
This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking…
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
Matthias Sperber, OndÅej Bojar, Barry Haddow +8
Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists o…