8 papers
SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills
Xuan Chen, Chengpeng Wang, Lu Yan +1
Agent skills extend LLM agents with reusable procedures, tools, and domain-specific workflows, but their safety depends on resolving dependencies among interacting instructions. We…
VERA: Variational Inference Framework for Jailbreaking Large Language Models
Anamika Lochab, Lu Yan, Patrick Pynadath +2
The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabilities in real-world settings. Without…
WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents
Lu Yan, Xuan Chen, Xiangyu Zhang
LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state. Existing instruction-follow…
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
Xuan Chen, Lu Yan, Ruqi Zhang +1
Large Language Model (LLM) agents increasingly act through external tools, making their safety contingent on tool-call workflows rather than text generation alone. While recent ben…
When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
Lu Yan, Xuan Chen, Xiangyu Zhang
Current coding-agent benchmarks usually pro- vide the full task specification upfront. Real research coding often does not: the intended system is progressively disclosed through i…
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
Xuan Chen, Shiwei Feng, Zikang Xiong +6
Assessing the safety of autonomous driving (AD) systems against security threats, particularly backdoor attacks, is a stepping stone for real-world deployment. However, existing wo…