2 papers
cs.CL2026
Noise Floor Audit for Agent Benchmarks
Yihang Chen, Pin Qian, Su Wang +4
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At tempera…
cs.AI2026
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification
Yihang Chen, Pin Qian, Su Wang +4
Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empir…