3 citations · 5 across the 6 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
On the Position Bias of On-Policy Distillation
Yan Xie, Sijie Zhu, Tiansheng Wen +2
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision from teachers. In the standard KL objective…
cs.LG2026
Scaling Attention via Feature Sparsity
Yan Xie, Tiansheng Wen, Tangda Huang +4
Scaling Transformers to ultra-long contexts is bottlenecked by the cost of self-attention. Existing methods reduce this cost along the sequence axis through local window…
cs.LG2026
Neural collapse in the orthoplex regime
James Alcala, Rayna Andreeva, Vladimir A. Kobzar +4
When training a neural network for classification, the feature vectors of the training set are known to collapse to the vertices of a regular simplex, provided the dimension of…