6 papers
The Kernel's Write: Application Read-Only Memory
Hui Sub Shim, Katherine Mohr, Philip Levis
Alongside power, DRAM has become a major limiting factor in datacenter growth. As DRAM's cost-per-bit has plateaued over the past decade, a class of emerging memory technologies, c…
Federation of Experts: Communication Efficient Distributed Inference for Large Language Models
Muhammad Shahir Abdurrahman, Chun Deng, Azalia Mirhoseini +1
Mixture of experts has emerged as the primary mechanism for making Large Language Models (LLMs) computationally efficient. However, in distributed settings, communicating token emb…
CvxCluster: Solving Large, Complex, Granular Resource Allocation Problems 100-1000x Faster
Obi Nnorom, Stephen Boyd, Philip Levis
Cluster resource allocation is a multidimensional search problem that finds the best allocation of tasks to servers. Because the search space grows exponentially, modern approaches…
The Future of Memory: Limits and Opportunities
Samuel Dayo, Shuhan Liu, Peijing Li +5
Memory latency, bandwidth, capacity, and energy increasingly limit performance. In this paper, we reconsider proposed system architectures that consist of huge (many-terabyte to pe…
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
Peijing Li, Muhammad Shahir Abdurraman, Rachel Cleaveland +6
Both SRAM and DRAM have stopped scaling: there is no technical roadmap to reduce their cost (per byte/GB). As a result, memory now dominates system cost. This paper argues for a pa…
GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition
Peijing Li, Matthew Hung, Yiming Tan +8
As AI workloads drive increasing memory requirements, domain-specific accelerators need higher-density on-chip memory beyond what current SRAM scaling trends can provide. Simultane…