1 paper · 1 filter
Raymond Bernard, Shaina Raza, Subhabrata Das +1
Despite the remarkable coherence of Large Language Models (LLMs), existing evaluation methods often suffer from fluency bias and rely heavily on multiple-choice formats, making it…