4 citations · 4 across the 2 of their papers we have counts for
4 papers · 1 filter
DiFR: Inference Verification Despite Nondeterminism
Adam Karvonen, Daniel Reuter, Roy Rinberg +3
As demand for LLM inference grows, it is becoming increasingly important that providers and their customers can verify that inference processes are performed correctly, without err…
Output Supervision Can Obfuscate the Chain of Thought
Jacob Drori, Luke Marks, Bryce Woodworth +2
OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep C…
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
Abir Harrasse, Philip Quirke, Clement Neo +3
Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQ…
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
Luke Marks, Alasdair Paren, David Krueger +1
Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that are not features of the input, limiting t…