1 paper · 1 filter
Xuetong Li, Gaofeng Liu
Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view cannot tell whether a model is…