1 paper · 1 filter
Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi +3
We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers…