From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
Huiyuan Tian, Bonan Xu, Shijian Li
The paper investigates how sparse mixture-of-experts language models route tokens to multiple experts, showing that expert subspaces overlap substantially yet routing still selects…
cs.CV2026
From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
Huiyuan Tian, Bonan Xu, Shijian Li
Feature-map knowledge distillation (KD) transfers internal representations well between comparably sized Vision Transformers (ViTs), but it often fails in compression. We revisit t…
cs.CV2025
SpectralKD: A Unified Framework for Interpreting and Distilling Vision Transformers via Spectral Analysis
Huiyuan Tian, Bonan Xu, Shijian Li +1
Knowledge Distillation (KD) has achieved widespread success in compressing large Vision Transformers (ViTs), but a unified theoretical framework for both ViTs and KD is still lacki…