Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus +23
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictiona…
cs.AI2025
Internal states before wait modulate reasoning patterns
Dmitrii Troitskii, Koyena Pal, Chris Wendler +2
Prior work has shown that a significant driver of performance in reasoning models is their ability to reason and self-correct. A distinctive marker in these reasoning traces is the…