Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
Zhonghao Zhan, Yefan Zhang, Krinos Li +1
Simulated evaluation is widely used to benchmark AI agents, yet how much evidence a simulated pass provides about physical deployment has not been systematically quantified. We pre…
cs.AI2026
How Adversarial Environments Mislead Agentic AI?
Zhonghao Zhan, Huichi Zhou, Zhenhao Li +3
Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack surface. Current evaluation…