collaborators

11 papers

cs.DL2026

Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems

Thales Sales Almeida, Giovana Kerche Bonás, Thiago Laitz +7

The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2…

cs.CL2026

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

Rodrigo Nogueira, Thales Sales Almeida, Giovana Kerche Bonás +7

Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, argui…

cs.CL2026

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6

Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to \emph{judge} such arguments -- weighing competin…

cs.CL2026

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6

Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering…

cs.CL2026

Measuring Opinion Bias and Sycophancy via LLM-based Persuasion

Rodrigo Nogueira, Giovana Kerche Bonás, Thales Sales Almeida +7

Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as agents, and used as a first sto…

cs.CL2026

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires +6

Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesi…