708 citations · 1.2k across the 25 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
Shaoyu Wang, Guangrong He, Geon-Woo Kim +2
Mixture-of-Experts (MoE) architectures offer the promise of larger model capacity without the prohibitive costs of fully dense designs. However, in real-world inference serving, lo…
cs.DC2022
Searching for Efficient Neural Architectures for On-Device ML on Edge TPUs
Berkin Akin, Suyog Gupta, Yun Long +6
On-device ML accelerators are becoming a standard in modern mobile system-on-chips (SoC). Neural architecture search (NAS) comes to the rescue for efficiently utilizing the high co…