2 papers
cs.AI2026
vToken: Token-Level Virtualization for Reclaimable KV Caches
Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen +4
Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allo…
cs.DC2026
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai +28
In-Network Collective (INC) acceleration holds immense potential for optimizing AI training and inference; however, its cross-layer nature has historically hindered investment and…