1 paper · 1 filter
Daniel Yoo, Adrians Skapars
Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generaliza…