1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Viacheslav Surkov, Chris Wendler, Antonio Mari +5
For large language models (LLMs), sparse autoencoders (SAEs) have been shown to decompose intermediate representations that often are not interpretable directly into sparse sums of…