1 paper · 1 filter
Parth Patil, Dhruv Kumar, Yash Sinha +1
Algebraic reasoning remains one of the most informative stress tests for large language models, yet current benchmarks provide no mechanism for attributing failure to a specific ca…