Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
Harshita Diddee, Gregory Yauney, Swabha Swayamdipta +1
Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality of benchmarks: a "poetry" benchma…
cs.CL2025
Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
Sayan Ghosh, Shahzaib Saqib Warraich, Dhruv Tarsadiya +2
Language models can be sampled multiple times to access the distribution underlying their responses, but existing methods cannot efficiently synthesize rich epistemic signals acros…
cs.CL2025
How Reliable is Language Model Micro-Benchmarking?
Gregory Yauney, Shahzaib Saqib Warraich, Swabha Swayamdipta
Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-b…