1 paper
A. Seza Doğruöz, Xixian Liao, Verena Blaschke +3
LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks (albeit mostly in English) due to shortcomings of conventional metrics and hig…