5 papers
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Michael Hardy, Anka Reuel, Lijin Zhang +6
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
Wenbo Zhang, Lijinghua Zhang, Liner Xiang +1
Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings remain unclear. Through contr…
Understanding Student Effort Using Response-Time Propensities During Problem Solving
Conrad Borchers, Lijin Zhang, Kexin Yang +2
Adaptive learning systems can produce substantial learning gains, yet many students engage for too brief or too superficial a period to benefit. A central obstacle is measuring eff…
Text Rationalization for Robust Causal Effect Estimation
Lijinghua Zhang, Hengrui Cai
Recent advances in natural language processing have enabled the increasing use of text data in causal inference, particularly for adjusting confounding factors in treatment effect…
Polytomous Explanatory Item Response Models for Item Discrimination: Assessing Negative-Framing Effects in Social-Emotional Learning Surveys
Joshua B. Gilbert, Lijin Zhang, Esther Ulitzsch +1
Modeling item parameters as a function of item characteristics has a long history but has generally focused on models for item location. Explanatory item response models for item d…