1 paper
Fumiaki Uehara, Koo Imai, Masato Tsutsumi +3
Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show that sparse autoencoder (SAE)…