Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages
Chanwoong Yoon, Jungsoo Park, Alan Ritter
Safety alignment helps models adhere to human values, but it often reduces response utility. We ask a critical but understudied question: Does safety alignment impose the cost equa…
cs.CL2026
Anticipatory Evaluation of Language Models
Jungsoo Park, Ethan Mendes, Gabriel Stanovsky +1
Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whethe…