14 citations · 25 across the 9 of their papers we have counts for
7 papers · 1 filter
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
María Grandury, Javier Aula-Blasco, Júlia Falcão +22
Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diver…
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
Raquel Ferrando, Javier Conde, Gonzalo Martínez +1
The computational and energy costs of Large Language Models (LLMs) have increased exponentially driven by the growing model sizes and the massive adoption of LLMs by hundreds of mi…
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
Javier Conde, Miguel González, María Grandury +3
The evaluation of LLMs has so far focused primarily on how well they can perform different tasks such as reasoning, question-answering, paraphrasing, or translating. For most of th…
Can ChatGPT Learn to Count Letters?
Javier Conde, Gonzalo Martínez, Pedro Reviriego +3
Large language models (LLMs) struggle on simple tasks such as counting the number of occurrences of a letter in a word. In this paper, we investigate if ChatGPT can learn to count…
Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models
Gonzalo Martínez, Javier Conde, Elena Merino-Gómez +4
Vocabulary tests, once a cornerstone of language modeling evaluation, have been largely overlooked in the current landscape of Large Language Models (LLMs) like Llama, Mistral, and…