2 papers
cs.CR2026
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
Kexin Chu, Zecheng Lin, Dawei Xiang +7
Global KV-cache sharing is an effective optimization for accelerating large language model (LLM) inference, yet it introduces an API-visible timing side channel that lets adversari…
cs.OS2025
XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing
Shushu Yi, Yuda An, Li Peng +11
Enterprise SSDs integrate numerous computing resources (e.g., ARM processor and onboard DRAM) to satisfy the ever-increasing performance requirements of I/O bursts. While these res…