1 paper
João N. Cardoso, Arlindo L. Oliveira, Bruno Martins
Understanding what features are encoded by learned directions in LLM activation space requires identifying inputs that strongly activate them. Feature visualization, which optimize…