3 papers
cs.AI2026
Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models
Sanjay Kariyappa, G. Edward Suh
Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the model should obey the highest-p…
cs.AI2026
Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
Sanjay Kariyappa, G. Edward Suh
Prompt injection attacks are a critical security vulnerability in large language models (LLMs), allowing attackers to hijack model behavior by injecting malicious instructions with…
cs.CL2025
Sequence-Level Leakage Risk of Training Data in Large Language Models
Trishita Tiwari, G. Edward Suh
This work quantifies the risk of training data leakage from LLMs (Large Language Models) using sequence-level probabilities. Computing extraction probabilities for individual seque…