activity
20222024
most citedEmbers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve

36 citations · 44 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20242 cited

Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning

Akshara Prabhakar, Thomas L. Griffiths, R. Thomas McCoy

Chain-of-Thought (CoT) prompting has been shown to enhance the multi-step reasoning capabilities of Large Language Models (LLMs). However, debates persist about whether LLMs exhibi…

cs.CL20242 cited

modeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models

Nathan A. Chi, Teodor Malchev, Riley Kong +5

We introduce modeLing, a novel benchmark of Linguistics Olympiad-style puzzles which tests few-shot reasoning in AI systems. Solving these puzzles necessitates inferring aspects of…

cs.CL202336 cited

Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve

R. Thomas McCoy, Shunyu Yao, Dan Friedman +2

The widespread adoption of large language models (LLMs) makes it important to recognize their strengths and limitations. We argue that in order to develop a holistic understanding…

cs.CL20233 cited

Modeling rapid language learning by distilling Bayesian priors into artificial neural networks

R. Thomas McCoy, Thomas L. Griffiths

Humans can learn languages from remarkably little experience. Developing computational models that explain this ability has been a major challenge in cognitive science. Bayesian mo…

cs.CL2022

Structural Biases for Improving Transformers on Translation into Morphologically Rich Languages

Paul Soulos, Sudha Rao, Caitlin Smith +9

Machine translation has seen rapid progress with the advent of Transformer-based models. These models have no explicit linguistic structure built into them, yet they may still impl…