3 papers
cs.CR2026
Combating Data Laundering in LLM Training
Muxing Li, Zesheng Ye, Sharon Li +1
Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprieta…
cs.LG2026
Distributional Statistics Restore Training Data Auditability in One-step Distilled Diffusion Models
Muxing Li, Zesheng Ye, Sharon Li +3
The proliferation of diffusion models trained on web-scale, provenance-uncertain image collections has made it essential, yet technically unresolved, to determine whether a model h…
cs.LG2026
Unlearning Evaluation through Subset Statistical Independence
Chenhao Zhang, Muxing Li, Feng Liu +2
Evaluating machine unlearning remains challenging, as existing methods typically require retraining reference models or performing membership inference attacks, both of which rely…