1 paper
Sanghee Park, Geewook Kim, Kee-Eung Kim
Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a decade (2014-2025) of Korean…