3 papers
cs.LG2025
Difficulties with Evaluating a Deception Detector for AIs
Lewis Smith, Bilal Chughtai, Neel Nanda
Building reliable deception detectors for AI systems -- methods that could predict when an AI system is being strategically deceptive without necessarily requiring behavioural evid…
cs.AI2025
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Nick Jiang, Xiaoqing Sun, Lisa Dunlap +2
Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data. Current metho…
cs.LG2025
Evaluating Sparse Autoencoders for Monosemantic Representation
Moghis Fereidouni, Muhammad Umair Haider, Peizhong Ju +1
A key barrier to interpreting large language models is polysemanticity, where neurons activate for multiple unrelated concepts. Sparse autoencoders (SAEs) have been proposed to mit…