1 paper
Xiangyu Wang, Jin Wu, Xiaoyu Li +2
Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are part…