2 citations · 4 across the 25 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Noise Floor Audit for Agent Benchmarks
Yihang Chen, Pin Qian, Su Wang +4
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At tempera…
cs.CL2026
Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict
Yihang Chen, Pin Qian, Su Wang +4
Retrieval-Augmented Generation (RAG) is usually evaluated by whether the final answer is correct. Under knowledge conflict, this hides a key question: did the model follow retrieve…