2 citations · 2 across the 9 of their papers we have counts for
22 papers · 1 filter
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde +2
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimize…
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
Samer Awad, Javier Conde, Carlos Arriaga +4
Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. While previous research has foc…
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
Tairan Fu, Javier Conde, Gonzalo MartÃnez +2
Multiple Choice Question (MCQ) tests are among the most used methods for evaluating large language models (LLMs). Besides checking the correctness of the selected answer, evaluatio…
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
Thomas Hikaru Clark, Carlos Arriaga, Javier Conde +2
Large Language Models (LLMs) have recently been shown to produce estimates of psycholinguistic norms, such as valence, arousal, or concreteness, for words and multiword expressions…
Large Language Models and Book Summarization: Reading or Remembering, Which Is Better?
Tairan Fu, Javier Conde, Pedro Reviriego +3
Summarization is a core task in Natural Language Processing (NLP). Recent advances in Large Language Models (LLMs) and the introduction of large context windows reaching millions o…
How does fine-tuning improve sensorimotor representations in large language models?
Minghua Wu, Javier Conde, Pedro Reviriego +1
Large Language Models (LLMs) exhibit a significant "embodiment gap", where their text-based representations fail to align with human sensorimotor experiences. This study systematic…