2 papers
cs.LG2026
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
Valeria Ruscio, Eli-Shaoul Khedouri, Keiran Thompson
Cross-entropy pretraining and preference alignment update the same transformer weights, but leave geometrically distinct traces. We characterise this asymmetry with a relative-subs…
cs.AI2026
The Phenomenology of Hallucinations
Valeria Ruscio, Keiran Thompson
We show that language models hallucinate not because they fail to detect uncertainty, but because of a failure to integrate it into output generation. Across architectures, uncerta…