activity
20242026
collaborators

6 papers

cs.CL2026

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

Sathvik Nair, Byung-Doh Oh

How predictable a word is can be quantified in two ways: using human responses to the cloze task or using probabilities from language models (LMs).When used as predictors of proces…

cs.CL2026

To model human linguistic prediction, make LLMs less superhuman

Byung-Doh Oh, Tal Linzen

When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs), which, like humans, make pred…

cs.CL2025

How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?

Christian Clark, Byung-Doh Oh, William Schuler

Contextual entropy is a psycholinguistic measure capturing the anticipated difficulty of processing a word just before it is encountered. Recent studies have tested for entropy-rel…

cs.CL2025

The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage

Byung-Doh Oh, Hongao Zhu, William Schuler

In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been sp…

cs.CL2025

The Impact of Token Granularity on the Predictive Power of Language Model Surprisal

Byung-Doh Oh, William Schuler

Word-by-word language model surprisal is often used to model the incremental processing of human readers, which raises questions about how various choices in language modeling infl…

cs.CL2024

Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities

Byung-Doh Oh, William Schuler

Predictions of word-by-word conditional probabilities from Transformer-based language models are often evaluated to model the incremental processing difficulty of human readers. In…