8 papers
OTRO: Oblivious Tokenization Path with Square-Root ORAM
Jonghyun Lee, Yongqin Wang, Rachit Rajat +3
The CPU-side large language model (LLM) tokenizer is a critical security gap in LLM serving through a confidential computing stack with CPU and GPU trusted execution environments (…
MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents
Haonan Li, Tianjun Sun, Yongqing Wang +1
Multi-server MCP agents create an information-flow control problem: faithful tool composition can turn individually benign read/write permissions into cross-boundary credential pro…
LRD-MPC: Efficient MPC Inference through Low-rank Decomposition
Tingting Tang, Yongqin Wang, Murali Annavaram
Secure Multi-party Computation (MPC) enables untrusted parties to jointly compute a function without revealing their inputs. Its application to machine learning (ML) has gained sig…
Differentially Private Retrieval-Augmented Generation
Tingting Tang, James Flemings, Yongqin Wang +1
Retrieval-augmented generation (RAG) is a widely used framework for reducing hallucinations in large language models (LLMs) on domain-specific tasks by retrieving relevant document…
Characterization of GPU TEE Overheads in Distributed Data Parallel ML Training
Jonghyun Lee, Yongqin Wang, Rachit Rajat +1
Confidential computing (CC) or trusted execution enclaves (TEEs) is now the most common approach to enable secure computing in the cloud. The recent introduction of GPU TEEs by NVI…
High-Throughput Secure Multiparty Computation with an Honest Majority in Various Network Settings
Christopher Harth-Kitzerow, Ajith Suresh, Yongqin Wang +3
In this work, we present novel protocols over rings for semi-honest secure three-party computation (3PC) and malicious four-party computation (4PC) with one corruption. While most…