38 citations · 71 across the 3 of their papers we have counts for
3 papers
cs.AR2024★ 1 cited
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
Yansong Xu, Dongxu Lyu, Zhenyu Li +6
Multi-scale deformable attention (MSDeformAttn) has emerged as a key mechanism in various vision tasks, demonstrating explicit superiority attributed to multi-scale grid-sampling.…
cs.DC2023★ 38 cited
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
Xiaonan Nie, Xupeng Miao, Zilong Wang +5
With the increasing data volume, there is a trend of using large-scale pre-trained models to store the knowledge into an enormous number of model parameters. The training of these…
cs.DC2022★ 32 cited
Tutel: Adaptive Mixture-of-Experts at Scale
Changho Hwang, Wei Cui, Yifan Xiong +12
Sparsely-gated mixture-of-experts (MoE) has been widely adopted to scale deep learning models to trillion-plus parameters with fixed computational cost. The algorithmic performance…