Publications (8)
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
Qinan Yu, Alexa Tartaglini, Peter Hase +2
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
Martha Lewis, Nihal V. Nayak, Peilin Yu +4
Large-scale neural network models combining text and images have made incredible progress in recent years. However, it remains an open question to what extent such models encode co…
Are Language Models Worse than Humans at Following Prompts? It's Complicated
Albert Webson, Alyssa Marie Loo, Qinan Yu +1
Prompts have been the center of progress in advancing language models' zero-shot and few-shot performance. However, recent work finds that models can perform surprisingly well when…
LLM Circuit Analyses Are Consistent Across Training and Scale
Curt Tigges, Michael Hanna, Qinan Yu +1
Most currently deployed large language models (LLMs) undergo continuous training or additional finetuning. By contrast, most research into LLMs' internal mechanisms focuses on mode…
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
Ruochen Zhang, Qinan Yu, Matianyu Zang +2
We employ new tools from mechanistic interpretability in order to ask whether the internal structure of large language models (LLMs) shows correspondence to the linguistic structur…
Grokking Group Multiplication with Cosets
Dashiell Stander, Qinan Yu, Honglu Fan +1
The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications. There have been many techniques developed to interpret deep n…