1 citations · 1 across the 8 of their papers we have counts for
Showing 2025 · cs.LGShow all
2 papers · 2 filters
cs.LG2025
ServerlessLoRA: Enabling Low-Latency Serverless Multi-LoRA Serving
Yifan Sui, Hao Wang, Hanfei Yu +5
Multi-LoRA (Low-Rank Adaptation) serving allows many specialized LLM variants to share the same base model by attaching lightweight adapters. This makes it attractive for serving l…
cs.LG2025
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
Hanfei Yu, Xingqi Cui, Hong Zhang +1
Large Language Models (LLMs) have gained immense success in revolutionizing various applications, including content generation, search and recommendation, and AI-assisted operation…