5 papers
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
José Pombal, Ricardo Rei, André F. T. Martins
LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by…
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
José Pombal, Nuno M. Guerreiro, Ricardo Rei +1
As language models improve and become capable of performing more complex tasks across modalities, evaluating them automatically becomes increasingly challenging. Developing strong…
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
Ricardo Rei, Nuno M. Guerreiro, José Pombal +4
Fine-tuning pretrained LLMs has been shown to be an effective strategy for reaching state-of-the-art performance on specific tasks like machine translation. However, this process o…
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
José Pombal, Nuno M. Guerreiro, Ricardo Rei +1
As automatic metrics become increasingly stronger and widely adopted, the risk of unintentionally "gaming the metric" during model development rises. This issue is caused by metric…
Analyzing Context Contributions in LLM-based Machine Translation
Emmanouil Zaranis, Nuno M. Guerreiro, André F. T. Martins
Large language models (LLMs) have achieved state-of-the-art performance in machine translation (MT) and demonstrated the ability to leverage in-context learning through few-shot ex…