Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Attention Sinks in Diffusion Transformers: A Causal Analysis
Fangzheng Wu, Brian Summa
Attention sinks -- tokens that receive disproportionate attention mass -- are assumed to be functionally important in autoregressive language models, but their role in diffusion tr…
cs.CV2026
Model-Centric Diagnostics: A Framework for Internal State Readouts
Fangzheng Wu, Brian Summa
We present a model-centric diagnostic framework that treats training state as a latent variable and unifies a family of internal readouts -- head-gradient norms, confidence, entrop…