1 paper
Alwin Jin, Sean M. Hendryx, Vaskar Nath
Current benchmarks that test LLMs on static, already-solved problems (e.g., math word problems) effectively demonstrated basic capability acquisition. The natural progression has b…