1 citations · 1 across the 1 of their papers we have counts for
1 paper
Keita Kiuchi, Yoshikazu Fujimoto, Hideyuki Gotō +5
Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional judgment. This observational fixed-benchmark…