cs.CL2026
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Yanshi Li, Xueru Bai, Shuman Liu +1
The paper introduces RepBench, a framework that aggregates thousands of benchmark datasets into a large set of probe texts to evaluate capability-aligned representations of large l…
#benchmark compilation#capability probing#representation learning#large language models