2 papers
cs.LG2026
Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation
Kumari Nishu, Han-Byul Kim, Santosh Chilkunda +13
Mixture-of-Experts (MoE) models are increasingly deployed alongside Speculative Decoding (SD) to accelerate inference, but combining the two is challenging. SD improves the inferen…
cs.LG2024
SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators
Rasoul Shafipour, David Harrison, Maxwell Horton +6
Large Language Models (LLMs) have transformed natural language processing, but face significant challenges in widespread deployment due to their high runtime cost. In this paper, w…