Showing cs.AIShow all
2 papers · 1 filter
cs.AI2024
On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
Alexander Bastounis, Paolo Campodonico, Mihaela van der Schaar +2
We introduce the Consistent Reasoning Paradox (CRP). Consistent reasoning, which lies at the core of human intelligence, is the ability to handle tasks that are equivalent, yet des…
cs.AI2024
Stealth edits to large language models
Oliver J. Sutton, Qinghua Zhou, Wei Wang +4
We reveal the theoretical foundations of techniques for editing large language models, and present new methods which can do so without requiring retraining. Our theoretical insight…