2 papers
cs.AI2026
Cascadia: Resident 975B MoE Inference on Eleven AI PCs
Tate Berenbaum, Matias Parij, Muthaiah Venkatachalam
Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coor…
cs.DC2026
Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure
Matias Parij, Pawan Paudel, Tate Berenbaum +1
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, sc…