3 papers
cs.CR2026
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
Guoci Chen, Xiurui Pan, Qiao Li +5
Deploying large language models (LLMs) as cloud services raises privacy concerns as inference may leak sensitive data. Fully Homomorphic Encryption (FHE) allows computation on encr…
cs.LG2025
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
Wenbo Wu, Qingyi Si, Xiurui Pan +2
While Key-Value (KV) cache succeeds in reducing redundant computations in auto-regressive models, it introduces significant memory overhead, limiting its practical deployment in lo…
cs.OS2025
XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing
Shushu Yi, Yuda An, Li Peng +11
Enterprise SSDs integrate numerous computing resources (e.g., ARM processor and onboard DRAM) to satisfy the ever-increasing performance requirements of I/O bursts. While these res…