2 papers
cs.LG2026
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
Muhammad Shahir Abdurrahman, Chun Deng, Azalia Mirhoseini +1
Mixture of experts has emerged as the primary mechanism for making Large Language Models (LLMs) computationally efficient. However, in distributed settings, communicating token emb…
cs.AR2025
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
Peijing Li, Muhammad Shahir Abdurraman, Rachel Cleaveland +6
Both SRAM and DRAM have stopped scaling: there is no technical roadmap to reduce their cost (per byte/GB). As a result, memory now dominates system cost. This paper argues for a pa…