Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch
Shaoyu Wang, Yizhuo Liang, Jaeyong Song +2
Mixture-of-Experts (MoE) architectures scale large language models (LLMs) to hundreds of billions of parameters. Serving a single MoE model requires multiple GPUs operating in para…
cs.DC2026
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
Jaeyong Song, Hongsun Jang, Changmin Shin +4
On-disk graph-based vector search (GVS) has become the dominant approach for serving large-scale vector databases at high recall, but prior systems struggle to sustain concurrent s…