1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.DC2025
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
Wentao Liu, Yuhao Hu, Ruiting Zhou +2
Mixture-of-Experts (MoE) has become a dominant architecture in large language models (LLMs) due to its ability to scale model capacity via sparse expert activation. Meanwhile, serv…
cs.CV2025
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
Youping Gu, Xiaolong Li, Yuhao Hu +2
Diffusion Transformers currently lead the field in high-quality video generation, but their slow iterative denoising process and prohibitive quadratic attention costs for long sequ…
cs.LG2024★ 1 cited
Ascend HiFloat8 Format for Deep Learning
Yuanyong Luo, Zhongxing Zhang, Richard Wu +16
This preliminary white paper proposes a novel 8-bit floating-point data format HiFloat8 (abbreviated as HiF8) for deep learning. HiF8 features tapered precision. For normal value e…