4 papers
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Rulin Shao, Akari Asai, Shannon Zejiang Shen +18
Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form…
Theoretical Analysis of Weak-to-Strong Generalization
Hunter Lang, David Sontag, Aravindan Vijayaraghavan
Strong student models can learn from weaker teachers: when trained on the predictions of a weaker model, a strong pretrained student can learn to correct the weak model's errors an…
Learning to Decode Collaboratively with Multiple Language Models
Shannon Zejiang Shen, Hunter Lang, Bailin Wang +2
We propose a method to teach multiple large language models (LLM) to collaborate by interleaving their generations at the token level. We model the decision of which LLM generates…
Towards Verifiable Text Generation with Symbolic References
Lucas Torroba Hennigen, Shannon Shen, Aniruddha Nrusimha +3
LLMs are vulnerable to hallucinations, and thus their outputs generally require laborious human verification for high-stakes applications. To this end, we propose symbolically grou…