1 paper · 1 filter
Florian A. D. Burnat, Brittany I. Davidson
Safety benchmarks are routinely treated as evidence about how a language model will behave once deployed, but this inference is fragile if behavior depends on whether a prompt look…