5 papers
-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing
Aoxi Liu, Yupeng Chen, James Oldfield +5
Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring for D-LLMs remains largely…
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
Panagiotis Koromilas, Andreas D. Demou, James Oldfield +2
Sparse autoencoders (SAEs) interpret neural network representations by decomposing activations into sparse combinations of dictionary atoms. However, SAEs assume features combine a…
fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery
Andreas D. Demou, Panagiotis Koromilas, James Oldfield +2
Many features in pretrained Transformers span multiple layers: they emerge through stages of inference, persist in the residual stream, or are built jointly by parallel MLPs. Cross…
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
James Oldfield, Philip Torr, Ioannis Patras +2
Monitoring large language models' (LLMs) activations is an effective way to detect harmful requests before they lead to unsafe outputs. However, traditional safety monitors often r…
Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
James Oldfield, Shawn Im, Sharon Li +3
Multilayer perceptrons (MLPs) are an integral part of large language models, yet their dense representations render them difficult to understand, edit, and steer. Recent methods le…