2 papers
cs.DC2026
Making MoE-based LLM Inference Resilient with Tarragon
Songyu Zhang, Aaron Tam, Myungjin Lee +2
Mixture-of-Experts (MoE) models are increasingly used to serve LLMs at scale, but failures become common as deployment scale grows. Existing systems exhibit poor failure resilience…
cs.NI2025
Palladium: A DPU-enabled Multi-Tenant Serverless Cloud over Zero-copy Multi-node RDMA Fabrics
Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan +3
Serverless computing promises enhanced resource efficiency and lower user costs, yet is burdened by a heavyweight, CPU-bound data plane. Prior efforts exploiting shared memory redu…