1 paper
Satchit Chatterji, Shihan Wang, Giovanni Sileno +1
Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what those fact…