1 paper · 1 filter
Alexander Pan, Lijie Chen, Jacob Steinhardt
Top-down transparency typically analyzes language model activations using probes with scalar or single-token outputs, limiting the range of behaviors that can be captured. To allev…