collaborators

15 papers

cs.DL2026

Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems

Thales Sales Almeida, Giovana Kerche Bonás, Thiago Laitz +7

The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2…

cs.CL2026

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz +2

Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand dee…

cs.CL2026

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT

Thiago Laitz, Thales Sales Almeida, João Guilherme Alves Santos +1

Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of…

cs.CL2026

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

Rodrigo Nogueira, Thales Sales Almeida, Giovana Kerche Bonás +7

Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, argui…

cs.CL2026

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6

Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to \emph{judge} such arguments -- weighing competin…

cs.CL2026

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6

Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering…