2 papers
cs.HC2025
Confirmation bias: A challenge for scalable oversight
Gabriel Recchia, Chatrik Singh Mangat, Jinu Nyachhyon +4
Scalable oversight protocols aim to empower evaluators to accurately verify AI models more capable than themselves. However, human evaluators are subject to biases that can lead to…
cs.AI2025
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
Gabriel Recchia, Chatrik Singh Mangat, Issac Li +1
As AI models tackle increasingly complex problems, ensuring reliable human oversight becomes more challenging due to the difficulty of verifying solutions. Approaches to scaling AI…