4 papers
Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages
Chanwoong Yoon, Jungsoo Park, Alan Ritter
Safety alignment helps models adhere to human values, but it often reduces response utility. We ask a critical but understudied question: Does safety alignment impose the cost equa…
Making Expert Reasoning Learnable with Self-Distillation
Ethan Mendes, Jungsoo Park, Alan Ritter
Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct solution to be reinforced or the existence o…
Anticipatory Evaluation of Language Models
Jungsoo Park, Ethan Mendes, Gabriel Stanovsky +1
Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whethe…
Language Models can Self-Improve at State-Value Estimation for Better Search
Ethan Mendes, Alan Ritter
Collecting ground-truth rewards or human demonstrations for multi-step reasoning tasks is often prohibitively expensive, particularly in interactive domains such as web tasks. We i…