4 papers
Position: Theory of Mind Benchmarks are Broken for Large Language Models
Matthew Riemer, Zahra Ashktorab, Djallel Bouneffouf +4
Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This…
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
Payel Das, Ching-Yun Ko, Sihui Dai +3
Large language models often expose their brittleness in reasoning tasks, especially while executing long chains of reasoning over context. We propose MemReasoner, a new and simple…
EpMAN: Episodic Memory AttentioN for Generalizing to Longer Contexts
Subhajit Chaudhury, Payel Das, Sarathkrishna Swaminathan +6
Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a signif…
Large Language Models can be Strong Self-Detoxifiers
Ching-Yun Ko, Pin-Yu Chen, Payel Das +6
Reducing the likelihood of generating harmful and toxic output is an essential task when aligning large language models (LLMs). Existing methods mainly rely on training an external…