2 papers
cs.DC2026
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
Hongyu Chen, Letian Ruan, Zilin Xu +6
LoRA enables efficient customization of LLMs and is widely used in multi-tenant and multi-task serving. However, emerging model architectures such as MoE significantly increase LoR…
cs.LG2025
Virtual Width Networks
Seed, Baisheng Li, Banggu Wu +115
We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN d…