1 paper · 1 filter
Simeng Han, Howard Dai, Stephen Xia +7
Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based…