Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Probing for Representation Manifolds in Superposition
Alexander Modell
This paper introduces the Manifold Probe, a supervised method for discovering representation manifolds in superposition. The method generalizes linear regression probes by learning…
cs.LG2025
The Origins of Representation Manifolds in Large Language Models
Alexander Modell, Patrick Rubin-Delanchy, Nick Whiteley
There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key el…