#benchmark validation
try —
2 papers match
cs.SE2026
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
Manyi Wang, Junjielong Xu, Pinjia He
The paper introduces PAIChecker, a multi‑agent system that automatically detects misalignments between pull requests and their linked issues in SWE‑bench‑style benchmarks, improvin…
#software engineering#large language models#benchmark validation#pr‑issue alignment
cs.AI2026
Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans
Jasmin Thelen, Oliver Wilhelm
The paper evaluates the ARC-AGI benchmark, originally designed for AI, as a measure of human fluid intelligence, finding good psychometric properties and a strong correlation with…
#fluid intelligence#rule induction#psychometrics#benchmark validation