Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
Dengke Han
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights…
cs.LG2024
Characterizing and Understanding HGNN Training on GPUs
Dengke Han, Mingyu Yan, Xiaochun Ye +1
Owing to their remarkable representation capabilities for heterogeneous graph data, Heterogeneous Graph Neural Networks (HGNNs) have been widely adopted in many critical real-world…