1 paper
Alex Flückiger, Chantal Amrhein, Tim Graf +5
As strong machine translation (MT) systems are increasingly based on large language models (LLMs), reliable quality benchmarking requires methods that capture their ability to leve…