6 papers
Continual Learning With Participation Privacy: An Auditable Buffering-Aggregation Recipe
T-H. Hubert Chan, Elaine Shi, Mengshi Zhao +1
Modern federated and streaming learning systems often release intermediate models, so privacy must hold for the full trajectory under adaptive interaction. Motivated by participati…
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Haowen Wang, Yaxin Du, Jian Yang +9
Mid-training has become an important stage in modern LLM development, using large-scale curated mixtures to strengthen capabilities before final post-training. Its data selection p…
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
Yijia Fang, Yiqing Feng, Bingyu Li +1
Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the adv…
MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
Sara Patel, Mingxun Zhou, Giulia Fanti
Generative search engines based on large language models (LLMs) are replacing traditional search, fundamentally changing how information providers are compensated. To sustain this…
PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
Jiahao Huo, Shuliang Liu, Bin Wang +5
Semantic-level watermarking (SWM) for large language models (LLMs) enhances watermarking robustness against text modifications and paraphrasing attacks by treating the sentence as…
Bifrost: A Much Simpler Secure Two-Party Data Join Protocol for Secure Data Analytics
Shuyu Chen, Mingxun Zhou, Haoyu Niu +2
Secure data join enables two parties with vertically distributed data to securely compute the joined table, allowing the parties to perform downstream Secure multi-party computatio…