4 papers
DIPBox: A Multi-scale Testing Framework for Tracking Dataset Regeneration
Tian Dong, Yan Meng, Shaofeng Li +5
Training datasets have tremendous proprietary value and are vulnerable to unauthorized copying. Existing defenses mainly focus on tracking individual data points, but pay little at…
Hunting Vulnerability Variants in AI Infra: Measurement and Reference-Driven Detection
Tian Dong, Yanjun Chen, Shoufeng Zhang +6
AI infra has become a shared execution layer for model training, deployment, and agent orchestration. Because many projects reimplement similar model-centric workflows, a vulnerabi…
Depth Gives a False Sense of Privacy: LLM Internal States Inversion
Tian Dong, Yan Meng, Shaofeng Li +3
Large Language Models (LLMs) are increasingly integrated into daily routines, yet they raise significant privacy and safety concerns. Recent research proposes collaborative inferen…
Model Inversion in Split Learning for Personalized LLMs: New Insights from Information Bottleneck Theory
Yunmeng Shu, Shaofeng Li, Tian Dong +2
Personalized Large Language Models (LLMs) have become increasingly prevalent, showcasing the impressive capabilities of models like GPT-4. This trend has also catalyzed extensive r…