4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.LG2025
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
HamidReza Imani, Jiaxin Peng, Peiman Mohseni +2
The deployment of mixture-of-experts (MoE) large language models (LLMs) presents significant challenges due to their high memory demands. These challenges become even more pronounc…
cs.DC2024★ 4 cited
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service
HamidReza Imani, Abdolah Amirany, Tarek El-Ghazawi
The increasing demand for deploying large Mixture-of-Experts (MoE) models in resource-constrained environments necessitates efficient approaches to address their high memory and co…