1 paper · 1 filter
Chunyang Li, Yilun Zheng, Xinting Huang +5
The paradigm of LLM-as-a-judge is emerging as a scalable and efficient alternative to human evaluation, demonstrating strong performance on well-defined tasks. However, its reliabi…