2 papers
cs.LG2025
Weight Ensembling Improves Reasoning in Language Models
Xingyu Dang, Christina Baek, Kaiyue Wen +2
We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. N…
cs.LG2025
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
Sachin Goyal, Christina Baek, J. Zico Kolter +1
A standard practice when using large language models is for users to supplement their instruction with an input context containing new information for the model to process. However…