1 citations · 1 across the 4 of their papers we have counts for
4 papers
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
Yuegui Huang, Zhiyuan Fang, Weiqi Luo +3
Despite the computational efficiency of MoE models, the excessive memory footprint and I/O overhead inherent in multi-expert architectures pose formidable challenges for real-time…
Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing
Junyi Wen, Junyuan Liang, Zicong Hong +3
Efficient state restoration in multi-turn conversations with large language models (LLMs) remains a critical challenge, primarily due to the overhead of recomputing or loading full…
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
Zhiyuan Fang, Yuegui Huang, Zicong Hong +5
Mixture of Experts (MoE), with its distinctive sparse structure, enables the scaling of language models up to trillions of parameters without significantly increasing computational…
Training and Serving System of Foundation Models: A Comprehensive Survey
Jiahang Zhou, Yanyu Chen, Zicong Hong +6
Foundation models (e.g., ChatGPT, DALL-E, PengCheng Mind, PanGu-) have demonstrated extraordinary performance in key technological areas, such as natural language processing and…