1 paper
Justin D. Norman, Michael U. Rivera, D. Alex Hughes
LLM-as-a-Judge has become the dominant evaluation paradigm for language models, but judge validation in practice relies on exact-match agreement, a metric that does not correct for…