1 paper · 1 filter
Adib Sakhawat, Tahsin Islam, Takia Farhin +3
The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequently conflate method variance,…