benchmarking 1conflict resolution 1constraint compliance 1instruction hierarchy 1llm robustness 1tool use 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CR2026
IH-Benchmark: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications
Conor McCauley, Zeliang Kan, Jason Martin
The paper introduces IH-Benchmark, a dataset for evaluating how large language models handle conflicting instructions across system‑user and user‑tool hierarchies, using a taxonomy…
cs.CR2025
LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
Sahar Abdelnabi, Aideen Fay, Ahmed Salem +22
Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models (LLMs) to distinguish between instructions and data in their inputs. Despite numerous def…