1 paper · 1 filter
Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk
Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation wit…