2 papers
cs.DC2026
PipeSwift: Revisiting Pipeline Parallelism for Large-Scale Completion-Oriented Agentic LLM Serving
Shiju Wang, Fei Ren, Fangcheng Fu +4
LLM agents execute long-horizon workflows where each model response determines the progress of subsequent tool interactions and environment transitions. Unlike chatbot serving, whe…
stat.ML2026
Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
Jia-Nan Wang, Zixun Huang, Kairui Li +1
We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability…