2 papers
cs.CR2026
Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale
Peichun Hua, Danyang Chen, Junan Zhang +5
Hosted retrieval-augmented generation (RAG) and semantic search allow users to query valuable provider-held corpora, raising two competing demands: to hide each query and chosen re…
cs.DC2026
CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?
Shuang Ma, Yuyi Li, Yihan Zhang +12
Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertis…