1 paper
Christabel Acquaye, Yi Ting Huang, Marine Carpuat +1
Standardized math assessments require expensive human pilot studies to establish the difficulty of test items. We investigate the predictive value of open-source large language mod…