Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
Eddie Yang, Dashun Wang
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep e…
cs.CL2024
Discovering influential text using convolutional neural networks
Megan Ayers, Luke Sanford, Margaret Roberts +1
Experimental methods for estimating the impacts of text on human evaluation have been widely used in the social sciences. However, researchers in experimental settings are usually…