papers

Publications (8)

cs.CL2026

Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning

Qinan Yu, Alexa Tartaglini, Peter Hase +2

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…

cs.CV2024

Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Martha Lewis, Nihal V. Nayak, Peilin Yu +4

Large-scale neural network models combining text and images have made incredible progress in recent years. However, it remains an open question to what extent such models encode co…

cs.CL2023

Are Language Models Worse than Humans at Following Prompts? It's Complicated

Albert Webson, Alyssa Marie Loo, Qinan Yu +1

Prompts have been the center of progress in advancing language models' zero-shot and few-shot performance. However, recent work finds that models can perform surprisingly well when…

cs.LG2024

LLM Circuit Analyses Are Consistent Across Training and Scale

Curt Tigges, Michael Hanna, Qinan Yu +1

Most currently deployed large language models (LLMs) undergo continuous training or additional finetuning. By contrast, most research into LLMs' internal mechanisms focuses on mode…

cs.CL2024

The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling

Ruochen Zhang, Qinan Yu, Matianyu Zang +2

We employ new tools from mechanistic interpretability in order to ask whether the internal structure of large language models (LLMs) shows correspondence to the linguistic structur…

cs.LG2024

Grokking Group Multiplication with Cosets

Dashiell Stander, Qinan Yu, Honglu Fan +1

The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications. There have been many techniques developed to interpret deep n…