2 papers
cs.DC2026
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving
Yifan Sui, Han Zhao, Rui Ma +6
LLM-powered agents execute tasks through a sequential loop of model generation and tool execution. Today's serving systems serialize this loop, leaving tool latency exposed on the…
cs.LG2025
ServerlessLoRA: Enabling Low-Latency Serverless Multi-LoRA Serving
Yifan Sui, Hao Wang, Hanfei Yu +5
Multi-LoRA (Low-Rank Adaptation) serving allows many specialized LLM variants to share the same base model by attaching lightweight adapters. This makes it attractive for serving l…