7 papers
Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation
Sang Truong, Yuheng Tu, Rylan Schaeffer +1
Scaling laws provide a fundamental framework for understanding the performance of Language Models (LMs), yet deriving them requires prohibitively expensive evaluations across thous…
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
Michael Hardy, Anka Reuel, Lijin Zhang +6
While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…
AI Evaluation Should Require Standardized Item-Level Data Releases
Han Jiang, Susu Zhang, Dongyao Zhu +6
This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…
Why Do Safety Guardrails Degrade Across Languages?
Max Zhang, Ameen Patel, Sang T. Truong +1
Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factor…
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Zeyu Tang, Sang T. Truong, Deonna Owens +4
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally un…
Fantastic Bugs and Where to Find Them in AI Benchmarks
Sang Truong, Yuheng Tu, Michael Hardy +8
Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of…