4 citations · 4 across the 1 of their papers we have counts for
5 papers
CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning
Saisai Xia, Wenhao Wang, Zihao Wang +4
Publicly available large pretrained models (i.e., backbones) and lightweight adapters for parameter-efficient fine-tuning (PEFT) have become standard components in modern machine l…
Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity
Guang Yan, Yuhui Zhang, Zimu Guo +6
With the growing use of large language models (LLMs) hosted on cloud platforms to offer inference services, privacy concerns about the potential leakage of sensitive information ar…
Trinity: A General Purpose FHE Accelerator
Xianglong Deng, Shengyu Fan, Zhicheng Hu +9
In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single…
Chameleon: An Efficient FHE Scheme Switching Acceleration on GPUs
Zhiwei Wang, Haoqi He, Lutan Zhao +4
Fully homomorphic encryption (FHE) enables direct computation on encrypted data, making it a crucial technology for privacy protection. However, FHE suffers from significant perfor…
The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems
Linke Song, Zixuan Pang, Wenhao Wang +7
The wide deployment of Large Language Models (LLMs) has given rise to strong demands for optimizing their inference performance. Today's techniques serving this purpose primarily f…