8 citations · 8 across the 4 of their papers we have counts for
1 paper · 1 filter
Charles Ye, Jasmine Cui
What is the most brute-force way to install interpretable, controllable features into a model's activations? Controlling how LLMs internally represent concepts typically requires s…