2 papers
cs.AI2026
Agentic Systems as Boosting Weak Reasoning Models
Varun Sunkaraneni, Pierfrancesco Beneventano, Riccardo Neumarker +2
Can a committee of weak reasoning-model calls reach the performance of much stronger models? We study verifier-backed committee search as inference-time boosting for reasoning lang…
cs.AI2026
The Generalized Turing Test: A Foundation for Comparing Intelligence
Daniel Mitropolsky, Susan S. Hong, Riccardo Neumarker +2
We introduce the Generalized Turing Test (GTT), a formal framework for comparing the capabilities of arbitrary agents via indistinguishability. For agents A and B, we define the Tu…