collaborators

5 papers

cs.CL2026

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

Tiago Teixeira, Ana Carolina Erthal, Juan Belieni +5

The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and benchmark datasets. However,…

cs.CL2026

Long-Context Generalization with Sparse Attention

Pavlo Vasylenko, Hugo Pitorro, André F. T. Martins +1

Transformer-based architectures traditionally employ softmax to compute attention weights, which produces dense distributions over all tokens in a sequence. While effective in many…

cs.CV2025

Movie Facts and Fibs (MF): A Benchmark for Long Movie Understanding

Emmanouil Zaranis, António Farinhas, Saul Santos +28

Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current be…

cs.CL2025

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi +17

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reprodu…

cs.CL2025

LegalBench.PT: A Benchmark for Portuguese Law

Beatriz Canaverde, Telmo Pessoa Pires, Leonor Melo Ribeiro +1

The recent application of LLMs to the legal field has spurred the creation of benchmarks across various jurisdictions and languages. However, no benchmark has yet been specifically…