1 paper
Chunyang Li, Yilun Zheng, Xinting Huang +5
The paradigm of LLM-as-a-judge is emerging as a scalable and efficient alternative to human evaluation, demonstrating strong performance on well-defined tasks. However, its reliabi…