1 citations · 1 across the 1 of their papers we have counts for
4 papers
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Iván Arcuschin, Jett Janiak, Robert Krzyzanowski +3
Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized…
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
Jett Janiak, Can Rager, James Dao +1
Prior work suggests that language models manage the limited bandwidth of the residual stream through a "memory management" mechanism, where certain attention heads and MLP layers c…
Characterizing stable regions in the residual stream of LLMs
Jett Janiak, Jacek Karwowski, Chatrik Singh Mangat +3
We identify stable regions in the residual stream of Transformers, where the model's output remains insensitive to small activation changes, but exhibits high sensitivity at region…
Evaluating Synthetic Activations composed of SAE Latents in GPT-2
Giorgi Giglemiani, Nora Petrova, Chatrik Singh Mangat +2
Sparse Auto-Encoders (SAEs) are commonly employed in mechanistic interpretability to decompose the residual stream into monosemantic SAE latents. Recent work demonstrates that pert…