17 citations · 17 across the 2 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu +10
Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instr…
cs.AI2024
Rule Based Rewards for Language Model Safety
Tong Mu, Alec Helyar, Johannes Heidecke +7
Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cas…