1 paper · 1 filter
Qihan Wang, Nicholas Tomlin, Michael Hu +2
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic m…