3 papers
cs.AI2026
Debate is efficient with your time
Jonah Brown-Cohen, Geoffrey Irving, Simon C. Marshall +3
AI safety via debate uses two competing models to help a human judge verify complex computational tasks. Previous work has established what problems debate can solve in principle,…
cs.CL2025
On the Brittleness of LLMs: A Journey around Set Membership
Lea Hergert, Gábor Berend, Mario Szegedy +2
Large language models (LLMs) achieve superhuman performance on complex reasoning tasks, yet often fail on much simpler problems, raising concerns about their reliability and interp…
cs.IT2024
Quantum Locally Testable Code with Constant Soundness
Andrew Cross, Zhiyang He, Anand Natarajan +2
In this paper, we present two constructions of quantum locally testable codes (QLTC) with constant soundness. In the first approach, we introduce an operation called check product,…