7 citations · 9 across the 11 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents
Lu Yan, Xuan Chen, Xiangyu Zhang
LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state. Existing instruction-follow…
cs.AI2024★ 1 cited
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
Guangyu Shen, Siyuan Cheng, Kaiyuan Zhang +6
Large Language Models (LLMs) have become prevalent across diverse sectors, transforming human life with their extraordinary reasoning and comprehension abilities. As they find incr…