most citedA suite of LMs comprehend puzzle statements as well as humans

1 citations · 1 across the 6 of their papers we have counts for

collaborators

12 papers

cs.CL2025

What Can String Probability Tell Us About Grammaticality?

Jennifer Hu, Ethan Gotlieb Wilcox, Siyuan Song +2

What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammatic…

cs.CL2025

Convergence and Divergence of Language Models under Different Random Seeds

Finlay Fehlauer, Kyle Mahowald, Tiago Pimentel

In this paper, we investigate the convergence of language models (LMs) trained under different random seeds, measuring convergence as the expected per-token Kullback--Leibler (KL)…

cs.AI2025

Privileged Self-Access Matters for Introspection in AI

Siyuan Song, Harvey Lederman, Jennifer Hu +1

Whether AI models can introspect is an increasingly important practical question. But there is no consensus on how introspection is to be defined. Beginning from a recently propose…

cs.CL2025

semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces

Jwalanthi Ranganathan, Rohan Jha, Kanishka Misra +1

We introduce semantic-features, an extensible, easy-to-use library based on Chronis et al. (2023) for studying contextualized word embeddings of LMs by projecting them into interpr…

cs.CL2025

Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs

William Sheffield, Kanishka Misra, Valentina Pyatkin +3

Discourse particles are crucial elements that subtly shape the meaning of text. These words, often polyfunctional, give rise to nuanced and often quite disparate semantic/discourse…

cs.CL20251 cited

A suite of LMs comprehend puzzle statements as well as humans

Adele E Goldberg, Supantho Rakshit, Jennifer Hu +1

Recent claims suggest that large language models (LMs) underperform humans in comprehending minimally complex English statements (Dentella et al., 2024). Here, we revisit those fin…