4 papers
Spiking the training data to correct for test set contamination
Johnny Tian-Zheng Wei, Jerry Li, Ameya Godbole +1
The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core proposal is to spike the training d…
SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion
Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3
Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full…
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
Mohammad Aflah Khan, Ameya Godbole, Johnny Tian-Zheng Wei +5
Understanding the relationship between training data and model behavior during pretraining is crucial, but existing workflows make this process cumbersome, fragmented, and often in…
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
Xinyue Cui, Johnny Tian-Zheng Wei, Swabha Swayamdipta +1
Data watermarking in language models injects traceable signals, such as specific token sequences or stylistic patterns, into copyrighted text, allowing copyright holders to track a…