1 paper · 1 filter
Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf +4
Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This…