1 paper · 1 filter
Charles Ye, Jasmine Cui
What is the most brute-force way to install interpretable, controllable features into a model's activations? Controlling how LLMs internally represent concepts typically requires s…