3 papers
cs.DC2026
EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs
Junming Zhang, Zhenzhe Zheng, Fan Wu +2
Mobile vendors and application developers increasingly deploy LLMs on smartphones for diverse prefill-only services. Yet current systems rely mainly on dense models whose regular c…
cs.CR2026
An Efficient and Privacy-Preserving Architecture for Cross-Institutional Collaborative RAG
Chenxin Mao, Shangyu Liu, Zhenzhe Zheng +3
Retrieval-Augmented Generation (RAG) empowers LLMs with external knowledge, making cross-institutional domain-specific knowledge base integration a highly promising deployment para…
cs.LG2025
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
Shangyu Liu, Zhenzhe Zheng, Xiaoyao Huang +3
Small language models (SLMs) support efficient deployments on resource-constrained edge devices, but their limited capacity compromises inference performance. Retrieval-augmented g…