From the 1 of 1 linked paper with an AI index.
1 paper
Conor McCauley, Zeliang Kan, Jason Martin
The paper introduces IH-Benchmark, a dataset for evaluating how large language models handle conflicting instructions across system‑user and user‑tool hierarchies, using a taxonomy…