2 papers
cs.LG2026
Attention-Only White-Box Transformer via LeJEPA-Based Self-Supervised Pretraining
Yang Bai, Linyuan Wang, Haoyang Jiang +3
Existing studies on self-supervised learning for white-box networks typically decouple the derivation of white-box networks via optimization algorithms from self-supervised learnin…
cs.CV2026
Interpretable and Sparse Linear Attention with Decoupled Membership-Subspace Modeling via MCR2 Objective
Tianyuan Liu, Libin Hou, Linyuan Wang +1
Maximal Coding Rate Reduction (MCR2)-driven white-box transformer, grounded in structured representation learning, unifies interpretability and efficiency, providing a reliable whi…