collaborators

7 papers

cs.LG2025

Pre-training under infinite compute

Konwoo Kim, Suhas Kotha, Percy Liang +1

Since compute grows much faster than web text available for language model pre-training, we ask how one should approach pre-training under fixed data and no compute constraints. We…

cs.LG2025

Learning to (Learn at Test Time): RNNs with Expressive Hidden States

Yu Sun, Xinhao Li, Karan Dalal +9

Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expre…

cs.CL2025

Auditing Prompt Caching in Language Model APIs

Chenchen Gu, Xiang Lisa Li, Rohith Kuditipudi +2

Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than non-cached prompts. These timing differences in…

cs.LG2025

Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Yann Dubois, Balázs Galambosi, Percy Liang +1

LLM-based auto-annotators have become a key component of the LLM development process due to their cost-effectiveness and scalability compared to human-based evaluation. However, th…

cs.CL2025

s1: Simple test-time scaling

Niklas Muennighoff, Zitong Yang, Weijia Shi +7

Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but…

cs.CL2025

AutoBencher: Towards Declarative Benchmark Construction

Xiang Lisa Li, Farzaan Kaiyom, Evan Zheran Liu +3

We present AutoBencher, a declarative framework for automatic benchmark construction, and use it to scalably discover novel insights and vulnerabilities of existing language models…