Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
(How) Do MLLMs Report Bistable Images Like Humans?
Ryota Takatsuki, Tomoki Doi, Amane Watahiki +2
Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans.…
cs.CV2025
Decoding Vision Transformers: the Diffusion Steering Lens
Ryota Takatsuki, Sonia Joseph, Ippei Fujisawa +1
Logit Lens is a widely adopted method for mechanistic interpretability of transformer-based language models, enabling the analysis of how internal representations evolve across lay…