Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Noise Floor Audit for Agent Benchmarks
Yihang Chen, Pin Qian, Su Wang +4
We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At tempera…
cs.CL2025
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
Zongxia Li, Yapei Chang, Yuhang Zhou +4
Evaluating open-ended long-form generation is challenging because it is hard to define what clearly separates good from bad outputs. Existing methods often miss key aspects like co…