artificial intelligence

Transcoders for Investigating Deception in Language Models

arXiv:2607.14791

summary

The paper uses per‑layer transcoders to build attribution graphs that reveal internal features linked to deceptive outputs in a Qwen3‑4B language model, showing how deception can be identified and monitored at the circuit level.

Abstract

Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour. In this paper, we investigate the use of transcoders to analyse deceptive behaviour in language models, a behaviour that poses a safety and security risk. Using a Qwen3-4B model with pre-trained transcoders, specifically per-layer transcoders (PLTs), we construct attribution graphs that capture feature activations and inter-feature dependencies, allowing circuit-level analysis of deception. Through feature steering and circuit analysis, we identified a dictionary of deception-related features and show that these features exert a stronger influence on deceptive outputs, as they produce predictable shifts between deceptive and non-deceptive responses. These findings suggest that deception emerges from internal model mechanisms and highlight the potential of transcoders for behavioural monitoring and early detection of security vulnerabilities related to malicious behaviours in language models.

Topics & keywords

#mechanistic interpretability#deception detection#language models#transcoders#circuit analysisper‑layer transcodersattribution graphsfeature steeringQwen3‑4Bdeceptive behavior
Transcoders for Investigating Deception in Language Models · wovepaper