2 papers
cs.LG2026
Gecko: Fast Private Inference via Secure Public Encoder Offloading
Cheng'an Wei, Kai Chen, Yue Zhao +2
Private inference protects both user inputs and server models during neural network inference, but existing solutions remain too slow for practical deployment. This motivates recen…
cs.CR2026
CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer
Yue Zhao, Yujia Gong, Ruigang Liang +4
The widespread deployment of large language models (LLMs) calls for post-hoc methods that can flexibly adapt models to evolving safety requirements. Meanwhile, the rapidly expandin…