1 paper · 1 filter
Rodrigo Guedes de Souza, Alison R. Panisson
Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., t…