collaborators

5 papers

cs.CL2024

Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation

Dávid Javorský, Dominik Macháček, Ondřej Bojar

Simultaneous speech translation (SST) can be evaluated on simulated online events where human evaluators watch subtitled videos and continuously express their satisfaction by press…

cs.CL2024

Multimodal Shannon Game with Images

Vilém Zouhar, Sunit Bhattacharya, Ondřej Bojar

The Shannon game has long been used as a thought experiment in linguistics and NLP, asking participants to guess the next letter in a sentence based on its preceding context. We ex…

cs.CL2024

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

Kartik Kartik, Sanjana Soni, Anoop Kunchukuttan +2

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This…

cs.CL2024

Understanding the role of FFNs in driving multilingual behaviour in LLMs

Sunit Bhattacharya, Ondřej Bojar

Multilingualism in Large Language Models (LLMs) is an yet under-explored area. In this paper, we conduct an in-depth analysis of the multilingual capabilities of a family of a Larg…

cs.CL2024

Quality and Quantity of Machine Translation References for Automatic Metrics

Vilém Zouhar, Ondřej Bojar

Automatic machine translation metrics typically rely on human translations to determine the quality of system translations. Common wisdom in the field dictates that the human refer…