2 papers
cs.CL2026
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Ying Wang, Zhen Jin, Jiexiong Xu +3
As augmented large language models (LLMs) with external tools become increasingly popular in web applications, improving augmented LLM inference serving efficiency and optimizing s…
cs.OS2025
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
Keyao Zhang, Yiquan Chen, Zhuo Hu +3
The accuracy of large language models (LLMs) improves with increasing model size, but increasing model complexity also poses significant challenges to training stability. Periodic…