1 paper · 1 filter
Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers +2
Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illumi…