1 citations · 3 across the 10 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2024
A General and Efficient Training for Transformer via Token Expansion
Wenxuan Huang, Yunhang Shen, Jiao Xie +5
The remarkable performance of Vision Transformers (ViTs) typically requires an extremely large training cost. Existing methods have attempted to accelerate the training of ViTs, ye…
cs.LG2024
Sinkhorn Distance Minimization for Knowledge Distillation
Xiao Cui, Yulei Qin, Yuting Gao +7
Knowledge distillation (KD) has been widely adopted to compress large language models (LLMs). Existing KD methods investigate various divergence measures including the Kullback-Lei…
cs.LG2023
SPD-DDPM: Denoising Diffusion Probabilistic Models in the Symmetric Positive Definite Space
Yunchen Li, Zhou Yu, Gaoqi He +4
Symmetric positive definite~(SPD) matrices have shown important value and applications in statistics and machine learning, such as FMRI analysis and traffic prediction. Previous wo…