Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
Can Jin, Jiakang Li, Rui Wu +3
As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong ge…
cs.AI2025
A Taxonomy of Transcendence
Natalie Abreu, Edwin Zhang, Eran Malach +1
Although language models are trained to mimic humans, the resulting systems display capabilities beyond the scope of any one person. To understand this phenomenon, we use a control…