From the 1 of 3 linked papers with an AI index.
1 citations · 1 across the 2 of their papers we have counts for
3 papers
Can We Trust Item Response Theory for AI Evaluation?
Han Jiang, Sunbeom Kwon, Jinwen Luo +2
The paper investigates how well item response theory (IRT) works for evaluating large language model benchmarks, highlighting challenges when benchmark data differ from traditional…
AI Evaluation Should Require Standardized Item-Level Data Releases
Han Jiang, Susu Zhang, Dongyao Zhu +6
This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…
Reducing Differential Item Functioning via Process Data
Ling Chen, Susu Zhang, Jingchen Liu
Testing fairness is a major concern in psychometric and educational research. A typical approach for ensuring testing fairness is through differential item functioning (DIF) analys…