From the 1 of 1 linked paper with an AI index.
1 paper
Darius Lim, Nathan Leow, Xin Wei Chia
The paper uses per‑layer transcoders to build attribution graphs that reveal internal features linked to deceptive outputs in a Qwen3‑4B language model, showing how deception can b…