2 citations · 4 across the 3 of their papers we have counts for
1 paper · 1 filter
Bulat Nutfullin, Vladimir Evgrafov, Dmitry Namiot
Prompt-based inference-time defenses for multimodal large language models (MLLMs) emit safety language into the very answers they defend, and a lexical keyword scorer whose refusal…