1 paper · 1 filter
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM saf…