1 paper
Junhao Luo, Ning Huang, Ziqi Sha +2
LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. In a controlled, solver-grounded sett…