3 papers
cs.DC2026
SpecGen: Accelerating Agentic Kernel Optimization with Speculative Generation
Jihu Guo, Sitian Lu, Tenghui Ma +4
Agentic kernel optimization automates manual GPU kernel tuning via iterative generation, validation, and profiling with reasoning LLMs, casting the optimization task as feedback-gu…
cs.DC2026
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
Wei Gao, Yuheng Zhao, Dilxat Muhtar +13
Agentic reinforcement learning (RL) is reshaping LLM post-training, but end-to-end training time is dominated by compute-intensive, multi-turn rollouts whose resource demand varies…
cs.LG2026
Lever: Speculative LLM Inference on Smartphones
Tuowei Wang, Fengzu Li, Yanfan Sun +2
Large language models (LLMs) are increasingly needed for interactive mobile applications, but high-quality models exceed the limited DRAM available on smartphones. Flash storage ca…