2 papers
cs.CR2026
Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened
Su Wang, Pin Qian, Yifan Lin +5
Self-improving AI agents are designed to learn from their mistakes. We show they can also hallucinate mistakes that never happened. We study this failure mode in automated harness…
cs.SE2026
When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
Su Wang, Pin Qian, Yihang Chen +6
LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agentic AI systems: whether indivi…