13 papers
Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems
Thales Sales Almeida, Giovana Kerche Bonás, Thiago Laitz +7
The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2…
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
Rodrigo Nogueira, Thales Sales Almeida, Giovana Kerche Bonás +7
Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, argui…
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6
Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to \emph{judge} such arguments -- weighing competin…
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
Roseval Malaquias Junior, Giovana Kerche Bonás, Thales Sales Almeida +6
Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scoring with multi-judge filtering…
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
Rodrigo Nogueira, Giovana Kerche Bonás, Thales Sales Almeida +7
Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as agents, and used as a first sto…
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
Thales Sales Almeida, Giovana Kerche Bonás, Ramon Pires +6
Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesi…