2 papers
cs.SE2026
Grounded Checklist Partial Credit for Agent Skill Trajectories
Suliu Qin, Lu Yin, Xilu Wang
Language-model agents increasingly tackle long-horizon tasks in interactive environments, yet their evaluation commonly relies on task-level success rates by reducing an entire exe…
cs.CR2026
AIRGuard: Guarding Agent Actions with Runtime Authority Control
Suliu Qin, Haomin Zhuang, Yujun Zhou +2
Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This ma…