Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Auto-Policy, not Auto-Skill: Compiled Agent Skills for the Physical World
Zhonghao Zhan, Hamed Haddadi
Self-evolving Skill harnesses (AutoSkills, Hermes Agent) generate more advisory orchestration automatically; their reported gains are efficiency, not safety. This misses the actual…
cs.AI2026
SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
Zhonghao Zhan, Yefan Zhang, Krinos Li +1
Simulated evaluation is widely used to benchmark AI agents, yet how much evidence a simulated pass provides about physical deployment has not been systematically quantified. We pre…
cs.AI2026
How Adversarial Environments Mislead Agentic AI?
Zhonghao Zhan, Huichi Zhou, Zhenhao Li +3
Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack surface. Current evaluation…