collaborators

5 papers

cs.CL2026

Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly

Yi-Chien Lin, William Schuler

There has been considerable interest in using surprisal from Transformer-based language models (LMs) as predictors of human sentence processing difficulty. Recent work has observed…

cs.CL2025

How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?

Christian Clark, Byung-Doh Oh, William Schuler

Contextual entropy is a psycholinguistic measure capturing the anticipated difficulty of processing a word just before it is encountered. Recent studies have tested for entropy-rel…

cs.CL2025

The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage

Byung-Doh Oh, Hongao Zhu, William Schuler

In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been sp…

cs.CL2025

The Impact of Token Granularity on the Predictive Power of Language Model Surprisal

Byung-Doh Oh, William Schuler

Word-by-word language model surprisal is often used to model the incremental processing of human readers, which raises questions about how various choices in language modeling infl…

cs.CL2025

Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled

Yi-Chien Lin, Hongao Zhu, William Schuler

The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive 'quality-power'…