1 paper · 1 filter
Yanbei Jiang, Amr Keleg, Ryandito Diandaru +4
While the real world is inherently stochastic, Large Language Models (LLMs) are predominantly evaluated on single-round inference against fixed ground truths. In this work, we shif…