1 paper · 1 filter
UniverseTBD, :, Kshitij Duraphe +3
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to v…