2 papers
cs.AI2026
In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?
Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi +2
Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LL…
cs.CV2025
Decoding Vision Transformers: the Diffusion Steering Lens
Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa +1
Logit Lens is a widely adopted method for mechanistic interpretability of transformer-based language models, enabling the analysis of how internal representations evolve across lay…