5 papers
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6
Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering…
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires +6
Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesi…
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
Giovana Kerche Bonás, Roseval Malaquias Junior, Marcos Piau +6
We introduce CAPITU, a benchmark for evaluating instruction-following capabilities of Large Language Models (LLMs) in Brazilian Portuguese. Unlike existing benchmarks that focus on…
The interplay between domain specialization and model size
Roseval Malaquias Junior, Ramon Pires, Thales Sales Almeida +3
Scaling laws for language models have often focused on finding the optimal model size and token count for training from scratch. However, achieving this optimal balance requires si…
Sabiá-3 Technical Report
Hugo Abonizio, Thales Sales Almeida, Thiago Laitz +4
This report presents Sabiá-3, our new flagship language model, and Sabiazinho-3, a more cost-effective sibling. The models were trained on a large brazilian-centric corpus. Evaluat…