4 papers · 1 filter
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Yang Xiao, Yusong Sun, Haoyi Wu +7
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends,…
RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
Yang Liu, Zhaokai Luo, Huayi Jin +6
As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limits GPU memory capacity, serv…
FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models
Bing Tian, Haikun Liu, Xiaocheng Zhong +5
Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly…
Akashic: A Low-Overhead LLM Inference Service with MemAttention
Yang Liu, Zhaokai Luo, Huayi Jin +7
Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every r…