5 papers
SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following
Beatriz Canaverde, Duarte M. Alves, José Pombal +2
In a conversation, a helpful assistant must reliably follow user directives, even as they refine, modify, or contradict earlier requests. Yet most instruction-following benchmarks…
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
Tiago Teixeira, Ana Carolina Erthal, Juan Belieni +5
The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and benchmark datasets. However,…
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
Afonso SimplÃcio, Gonçalo Vinagre, Miguel Moura Ramos +19
Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translat…
Movie Facts and Fibs (MF): A Benchmark for Long Movie Understanding
Emmanouil Zaranis, António Farinhas, Saul Santos +28
Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current be…
LegalBench.PT: A Benchmark for Portuguese Law
Beatriz Canaverde, Telmo Pessoa Pires, Leonor Melo Ribeiro +1
The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically…