5 papers
Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs
Charles Westphal, Timothy Douglas, Keivan Navaie +2
Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltration risk that is difficult to d…
Towards Worst-Case Guarantees with Scale-Aware Interpretability
Lauren Greenspan, David Berman, Aryeh Brill +9
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…
Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…
Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
Charles Westphal, Keivan Navaie, Fernando E. Rosas
Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied on trivially recoverable encodings. We f…
Explosive neural networks via higher-order interactions in curved statistical manifolds
Miguel Aguilera, Pablo A. Morales, Fernando E. Rosas +1
Higher-order interactions underlie complex phenomena in systems such as biological and artificial neural networks, but their study is challenging due to the scarcity of tractable m…