4 papers
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
Thales Sales Almeida, João Guilherme Alves Santos, Thiago Laitz +1
Large language models (LLMs) are increasingly deployed as task-oriented agents, where success depends on their ability to generate accurate function calls under realistic, multilin…
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
Thales Sales Almeida, Giovana Kerche Bonás, João Guilherme Alves Santos
Large Language Models (LLMs) exhibit significant performance variations depending on the linguistic and cultural context in which they are applied. This disparity signals the neces…
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
João Guilherme Alves Santos, Giovana Kerche Bonás, Thales Sales Almeida
With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We pr…
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
Thales Sales Almeida, Giovana Kerche Bonás, João Guilherme Alves Santos +2
As the knowledge landscape evolves and large language models (LLMs) become increasingly widespread, there is a growing need to keep these models updated with current events. While…