activity
20242026
most citedAI Evaluation Should Require Standardized Item-Level Data Releases

1 citations · 2 across the 8 of their papers we have counts for

collaborators

12 papers

cs.CY2026

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Meera Desai, Sang T. Truong, Hanna Wallach +8

Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reason…

cs.CL2026

A Dataset for Modeling Iterative Problem-Solving

Fagun Patel, Sang T. Truong, Duc Q. Nguyen +4

Solving problems through repeated attempts is a sequential modeling task: at each step, the solver receives feedback and decides how to revise their solutions. Predicting whether p…

cs.LG2026

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

Sang Truong, Yuheng Tu, Rylan Schaeffer +1

Scaling laws provide a fundamental framework for understanding the performance of Language Models (LMs), yet deriving them requires prohibitively expensive evaluations across thous…

cs.AI2026

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

Michael Hardy, Anka Reuel, Lijin Zhang +6

While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…

cs.AI20261 cited

AI Evaluation Should Require Standardized Item-Level Data Releases

Han Jiang, Susu Zhang, Dongyao Zhu +6

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…

cs.CL2026

Why Do Safety Guardrails Degrade Across Languages?

Max Zhang, Ameen Patel, Sang T. Truong +1

Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factor…