25 papers
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…
Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
Tairan Fu, Javier Conde, Carlos Arriaga +5
Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized consensus even for open questio…
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
Samer Awad, Javier Conde, Carlos Arriaga +4
Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. While previous research has foc…
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
Tairan Fu, Javier Conde, Gonzalo MartÃnez +2
Multiple Choice Question (MCQ) tests are among the most used methods for evaluating large language models (LLMs). Besides checking the correctness of the selected answer, evaluatio…
Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test
Tairan Fu, Francisco Javier Santos-MartÃn, Javier Conde +2
The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. W…
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
Thomas Hikaru Clark, Carlos Arriaga, Javier Conde +2
Large Language Models (LLMs) have recently been shown to produce estimates of psycholinguistic norms, such as valence, arousal, or concreteness, for words and multiword expressions…