activity
20242026
collaborators

5 papers

cs.CL2026

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6

Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering…

cs.CL2026

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires +6

Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesi…

cs.CL2026

CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context

Giovana Kerche Bonás, Roseval Malaquias Junior, Marcos Piau +6

We introduce CAPITU, a benchmark for evaluating instruction-following capabilities of Large Language Models (LLMs) in Brazilian Portuguese. Unlike existing benchmarks that focus on…

cs.CL2025

The interplay between domain specialization and model size

Roseval Malaquias Junior, Ramon Pires, Thales Sales Almeida +3

Scaling laws for language models have often focused on finding the optimal model size and token count for training from scratch. However, achieving this optimal balance requires si…

cs.CL2024

Sabiá-3 Technical Report

Hugo Abonizio, Thales Sales Almeida, Thiago Laitz +4

This report presents Sabiá-3, our new flagship language model, and Sabiazinho-3, a more cost-effective sibling. The models were trained on a large brazilian-centric corpus. Evaluat…