1 paper · 1 filter
Hamid Osooli, Kareema Batool, Rick Gentry +3
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's bl…