10 citations · 10 across the 1 of their papers we have counts for
1 paper
Hongru Liang, Huaqing Li
Human evaluation is becoming a necessity to test the performance of Chatbots. However, off-the-shelf settings suffer the severe reliability and replication issues partly because of…