Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
Isaac Llorente-Saguer
Aligned language models refuse harmful instructions, but the representations through which they recognise such instructions are less well characterised than the behaviours they pro…
cs.LG2026
The Geometry of Harmful Intent: Training-Free Anomaly Detection via Angular Deviation in LLM Residual Streams
Isaac Llorente-Saguer
We present LatentBiopsy, a training-free method for detecting harmful prompts by analysing the geometry of residual-stream activations in large language models. Given 200 safe norm…