1 citations · 1 across the 4 of their papers we have counts for
4 papers
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
Yuegui Huang, Zhiyuan Fang, Weiqi Luo +3
Despite the computational efficiency of MoE models, the excessive memory footprint and I/O overhead inherent in multi-expert architectures pose formidable challenges for real-time…
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
Zhiyuan Fang, Yuegui Huang, Zicong Hong +5
Mixture of Experts (MoE), with its distinctive sparse structure, enables the scaling of language models up to trillions of parameters without significantly increasing computational…
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
Zhiyuan Fang, Zicong Hong, Yuegui Huang +5
Large Language Models (LLMs) have demonstrated impressive performance across various tasks, and their application in edge scenarios has attracted significant attention. However, sp…
SmartOracle: Generating Smart Contract Oracle via Fine-Grained Invariant Detection
Jianzhong Su, Jiachi Chen, Zhiyuan Fang +3
As decentralized applications (DApps) proliferate, the increased complexity and usage of smart contracts have heightened their susceptibility to security incidents and financial lo…