3 papers
cs.CR2026
Differentially Private and Communication Efficient Large Language Model Split Inference via Stochastic Quantization and Soft Prompt
Yujie Gu, Richeng Jin, Xiaoyu Ji +2
Large Language Models (LLMs) have achieved remarkable performance and received significant research interest. The enormous computational demands, however, hinder the local deployme…
cs.CR2025
CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning
Saisai Xia, Wenhao Wang, Zihao Wang +4
Publicly available large pretrained models (i.e., backbones) and lightweight adapters for parameter-efficient fine-tuning (PEFT) have become standard components in modern machine l…
cs.CR2025
The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems
Linke Song, Zixuan Pang, Wenhao Wang +7
The wide deployment of Large Language Models (LLMs) has given rise to strong demands for optimizing their inference performance. Today's techniques serving this purpose primarily f…