3 papers
cs.CR2026
Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs
Charles Westphal, Timothy Douglas, Keivan Navaie +2
Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltration risk that is difficult to d…
hep-th2026
Towards Worst-Case Guarantees with Scale-Aware Interpretability
Lauren Greenspan, David Berman, Aryeh Brill +9
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…
cs.LG2026
Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…