7 papers
Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems
Thales Sales Almeida, Giovana Kerche Bonás, Thiago Laitz +7
The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer Society since 2…
BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams
João Guilherme Alves Santos, Giovana Kerche Bonás, Thiago Laitz +2
Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand dee…
moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT
Thiago Laitz, Thales Sales Almeida, João Guilherme Alves Santos +1
Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of…
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
Thales Sales Almeida, João Guilherme Alves Santos, Thiago Laitz +1
Large language models (LLMs) are increasingly deployed as task-oriented agents, where success depends on their ability to generate accurate function calls under realistic, multilin…
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
Thales Sales Almeida, Giovana Kerche Bonás, João Guilherme Alves Santos
Large Language Models (LLMs) exhibit significant performance variations depending on the linguistic and cultural context in which they are applied. This disparity signals the neces…
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
João Guilherme Alves Santos, Giovana Kerche Bonás, Thales Sales Almeida
With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We pr…