most citedComet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity

4 citations · 4 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CR2025

CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning

Saisai Xia, Wenhao Wang, Zihao Wang +4

Publicly available large pretrained models (i.e., backbones) and lightweight adapters for parameter-efficient fine-tuning (PEFT) have become standard components in modern machine l…

cs.CR20254 cited

Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity

Guang Yan, Yuhui Zhang, Zimu Guo +6

With the growing use of large language models (LLMs) hosted on cloud platforms to offer inference services, privacy concerns about the potential leakage of sensitive information ar…

cs.AR2024

Trinity: A General Purpose FHE Accelerator

Xianglong Deng, Shengyu Fan, Zhicheng Hu +9

In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single…

cs.CR2024

Chameleon: An Efficient FHE Scheme Switching Acceleration on GPUs

Zhiwei Wang, Haoqi He, Lutan Zhao +4

Fully homomorphic encryption (FHE) enables direct computation on encrypted data, making it a crucial technology for privacy protection. However, FHE suffers from significant perfor…

cs.CR2024

The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems

Linke Song, Zixuan Pang, Wenhao Wang +7

The wide deployment of Large Language Models (LLMs) has given rise to strong demands for optimizing their inference performance. Today's techniques serving this purpose primarily f…