papers

Publications (17)

cs.AI2025

Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Violet Xiang, Charlie Snell, Kanishk Gandhi +11

We propose a novel framework, Meta Chain-of-Thought (Meta-CoT), which extends traditional Chain-of-Thought (CoT) by explicitly modeling the underlying reasoning required to arrive…

cs.CL2023

The False Promise of Imitating Proprietary LLMs

Arnav Gudibande, Eric Wallace, Charlie Snell +5

An emerging method to cheaply improve a weaker language model is to finetune it on outputs from a stronger model, such as a proprietary system like ChatGPT (e.g., Alpaca, Self-Inst…

cs.LG2025

Value-Based Deep RL Scales Predictably

Oleh Rybkin, Michal Nauman, Preston Fu +4

Scaling data and compute is critical to the success of modern ML. However, scaling demands predictability: we want methods to not only perform well with more compute or data, but a…

cs.AI2025

Sleep-time Compute: Beyond Inference Scaling at Test-time

Kevin Lin, Charlie Snell, Yu Wang +4

Scaling test-time compute has emerged as a key ingredient for enabling large language models (LLMs) to solve difficult problems, but comes with high latency and inference cost. We…

cs.CL2023

Non-Programmers Can Label Programs Indirectly via Active Examples: A Case Study with Text-to-SQL

Ruiqi Zhong, Charlie Snell, Dan Klein +1

Can non-programmers annotate natural language utterances with complex programs that represent their meaning? We introduce APEL, a framework in which non-programmers select among ca…

cs.LG2024

Predicting Emergent Capabilities by Finetuning

Charlie Snell, Eric Wallace, Dan Klein +1

A fundamental open challenge in modern LLM scaling is the lack of understanding around emergent capabilities. In particular, language model pretraining loss is known to be highly p…

cs.LG2024

Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Charlie Snell, Jaehoon Lee, Kelvin Xu +1

Enabling LLMs to improve their outputs by using more test-time computation is a critical step towards building generally self-improving agents that can operate on open-ended natura…

cs.LG2025

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Amrith Setlur, Matthew Y. R. Yang, Charlie Snell +5

Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.…

cs.AI2025

Reasoning Models Can Be Effective Without Thinking

Wenjie Ma, Jingxuan He, Charlie Snell +3

Recent LLMs have significantly improved reasoning capabilities, primarily by including an explicit, lengthy Thinking process as part of generation. In this paper, we question wheth…

cs.CL2023

Offline RL for Natural Language Generation with Implicit Language Q Learning

Charlie Snell, Ilya Kostrikov, Yi Su +2

Large language models distill broad knowledge from text corpora. However, they can be inconsistent when it comes to completing user specified tasks. This issue can be addressed by…

cs.CL2022

Context-Aware Language Modeling for Goal-Oriented Dialogue Systems

Charlie Snell, Mengjiao Yang, Justin Fu +2

Goal-oriented dialogue systems face a trade-off between fluent language generation and task-specific control. While supervised learning with large language models is capable of pro…

cs.AI2025

Learning Adaptive Parallel Reasoning with Language Models

Jiayi Pan, Xiuyu Li, Long Lian +6

Scaling inference-time computation has substantially improved the reasoning capabilities of language models. However, existing methods have significant limitations: serialized chai…

cs.CL2021

Approximating How Single Head Attention Learns

Charlie Snell, Ruiqi Zhong, Dan Klein +1

Why do models often attend to salient words, and how does this evolve throughout training? We approximate model training as a two stage process: early on in training when the atten…

cs.CL2023

LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Marwa Abdulhai, Isadora White, Charlie Snell +5

Large language models (LLMs) provide excellent text-generation capabilities, but standard prompting and generation methods generally do not lead to intentional or goal-directed age…

cs.CL2022

Describing Differences between Text Distributions with Natural Language

Ruiqi Zhong, Charlie Snell, Dan Klein +1

How do two distributions of texts differ? Humans are slow at answering this, since discovering patterns might require tediously reading through hundreds of samples. We propose to a…

cs.SE2026

Composer 2 Technical Report

Cursor Research, :, Aaron Chan +53

Composer 2 is a specialized model designed for agentic software engineering. The model demonstrates strong long-term planning and coding intelligence while maintaining the ability…

cs.CL2022

Learning by Distilling Context

Charlie Snell, Dan Klein, Ruiqi Zhong

Language models significantly benefit from context tokens, such as prompts or scratchpads. They perform better when prompted with informative instructions, and they acquire new rea…