Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models
Sanjay Kariyappa, G. Edward Suh
Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the model should obey the highest-p…
cs.AI2026
Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
Sanjay Kariyappa, G. Edward Suh
Prompt injection attacks are a critical security vulnerability in large language models (LLMs), allowing attackers to hijack model behavior by injecting malicious instructions with…