Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AI Evaluation Should Require Standardized Item-Level Data Releases
Han Jiang, Susu Zhang, Dongyao Zhu +6
This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…
cs.AI2025
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
Jing Yao, Xiaoyuan Yi, Shitong Duan +8
As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applicati…