1 paper · 1 filter
Zhichao Fan, Yanhang Li, Zexin Zhuang +2
Trust-benchmark scores reported on a chat-LLM release line are often carried across several checkpoints of the same line, as if the underlying model had not shifted between release…