20 papers
Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems
Thales Sales Almeida, Giovana Kerche Bonás, Thiago Laitz +7
The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2…
BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz +2
Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand dee…
moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT
Thiago Laitz, Thales Sales Almeida, João Guilherme Alves Santos +1
Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of…
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
Rodrigo Nogueira, Thales Sales Almeida, Giovana Kerche Bonás +7
Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, argui…
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6
Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to \emph{judge} such arguments -- weighing competin…
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6
Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering…