4 papers
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
José Pombal, Ricardo Rei, André F. T. Martins
LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB): they tend to favor outputs produced by…
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
José Pombal, Nuno M. Guerreiro, Ricardo Rei +1
As language models improve and become capable of performing more complex tasks across modalities, evaluating them automatically becomes increasingly challenging. Developing strong…
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
Ricardo Rei, Nuno M. Guerreiro, José Pombal +4
Fine-tuning pretrained LLMs has been shown to be an effective strategy for reaching state-of-the-art performance on specific tasks like machine translation. However, this process o…
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
José Pombal, Nuno M. Guerreiro, Ricardo Rei +1
As automatic metrics become increasingly stronger and widely adopted, the risk of unintentionally "gaming the metric" during model development rises. This issue is caused by metric…