9 citations · 17 across the 9 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Mechanisms of Introspective Awareness
Uzay Macar, Li Yang, Atticus Wang +3
Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introsp…
cs.LG2025★ 9 cited
Open Problems in Mechanistic Interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson +26
Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…