Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy +7
Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI systems. We tested this assumpti…
cs.AI2024
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi +6
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not…