1 paper · 1 filter
Chia-Yu Hsu, Shubhanshu Shekhar
We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct…