1 paper
Yuhan Ma, Yong Li, Stefan Schmid
Two-server secure inference allows a client to query a hosted large language model (LLM) without revealing prompts or embeddings. Recent GPU systems based on function secret sharin…