activity
20232026
most citedUnderstanding the Impact of Artificial Intelligence in Academic Writing: Metadata to the Rescue

14 citations · 25 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2

Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…

cs.CL2025

La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America

María Grandury, Javier Aula-Blasco, Júlia Falcão +22

Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diver…

cs.CL2025

Is There a Case for Conversation Optimized Tokenizers in Large Language Models?

Raquel Ferrando, Javier Conde, Gonzalo Martínez +1

The computational and energy costs of Large Language Models (LLMs) have increased exponentially driven by the growing model sizes and the massive adoption of LLMs by hundreds of mi…

cs.CL2025★ 1 cited

Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans

Javier Conde, Miguel González, María Grandury +3

The evaluation of LLMs has so far focused primarily on how well they can perform different tasks such as reasoning, question-answering, paraphrasing, or translating. For most of th…

cs.CL2025★ 6 cited

Can ChatGPT Learn to Count Letters?

Javier Conde, Gonzalo Martínez, Pedro Reviriego +3

Large language models (LLMs) struggle on simple tasks such as counting the number of occurrences of a letter in a word. In this paper, we investigate if ChatGPT can learn to count…

cs.CL2023

Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models

Gonzalo Martínez, Javier Conde, Elena Merino-Gómez +4

Vocabulary tests, once a cornerstone of language modeling evaluation, have been largely overlooked in the current landscape of Large Language Models (LLMs) like Llama, Mistral, and…