collaborators

6 papers

cs.AR2026

The Kernel's Write: Application Read-Only Memory

Hui Sub Shim, Katherine Mohr, Philip Levis

Alongside power, DRAM has become a major limiting factor in datacenter growth. As DRAM's cost-per-bit has plateaued over the past decade, a class of emerging memory technologies, c…

cs.LG2026

Federation of Experts: Communication Efficient Distributed Inference for Large Language Models

Muhammad Shahir Abdurrahman, Chun Deng, Azalia Mirhoseini +1

Mixture of experts has emerged as the primary mechanism for making Large Language Models (LLMs) computationally efficient. However, in distributed settings, communicating token emb…

cs.DC2026

CvxCluster: Solving Large, Complex, Granular Resource Allocation Problems 100-1000x Faster

Obi Nnorom, Stephen Boyd, Philip Levis

Cluster resource allocation is a multidimensional search problem that finds the best allocation of tasks to servers. Because the search space grows exponentially, modern approaches…

cs.AR2025

The Future of Memory: Limits and Opportunities

Samuel Dayo, Shuhan Liu, Peijing Li +5

Memory latency, bandwidth, capacity, and energy increasingly limit performance. In this paper, we reconsider proposed system architectures that consist of huge (many-terabyte to pe…

cs.AR2025

Towards Memory Specialization: A Case for Long-Term and Short-Term RAM

Peijing Li, Muhammad Shahir Abdurraman, Rachel Cleaveland +6

Both SRAM and DRAM have stopped scaling: there is no technical roadmap to reduce their cost (per byte/GB). As a result, memory now dominates system cost. This paper argues for a pa…

cs.AR2025

GainSight: A Unified Framework for Data Lifetime Profiling and Heterogeneous Memory Composition

Peijing Li, Matthew Hung, Yiming Tan +8

As AI workloads drive increasing memory requirements, domain-specific accelerators need higher-density on-chip memory beyond what current SRAM scaling trends can provide. Simultane…