1 paper
Qiyao Wei, Edward Morrell, Lea Goetz +1
Evaluating the open-form textual responses generated by Large Language Models (LLMs) typically requires measuring the semantic similarity of the response to a (human generated) ref…