collaborators

9 papers

stat.ME2026

Spiking the training data to correct for test set contamination

Johnny Tian-Zheng Wei, Jerry Li, Ameya Godbole +1

The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core proposal is to spike the training d…

cs.LG2026

SHRED: Retain-Set-Free Unlearning via Self-Distillation with Logit Demotion

Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei +3

Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full…

cs.CL2026

Psychological Steering of Large Language Models

Leonardo Blas, Robin Jia, Emilio Ferrara

Large language models (LLMs) emulate a consistent human-like behavior that can be shaped through activation-level interventions. This paradigm is converging on additive residual-st…

cs.CL2025

Hubble: a Model Suite to Advance the Study of LLM Memorization

Johnny Tian-Zheng Wei, Ameya Godbole, Mohammad Aflah Khan +7

We present Hubble, a suite of fully open-source large language models (LLMs) for the scientific study of LLM memorization. Hubble models come in standard and perturbed variants: st…

cs.CL2025

Teaching Models to Understand (but not Generate) High-risk Data

Ryan Wang, Matthew Finlayson, Luca Soldaini +2

Language model developers typically filter out high-risk content -- such as toxic or copyrighted text -- from their pre-training data to prevent models from generating similar outp…

cs.CL2025

TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability

Mohammad Aflah Khan, Ameya Godbole, Johnny Tian-Zheng Wei +5

Understanding the relationship between training data and model behavior during pretraining is crucial, but existing workflows make this process cumbersome, fragmented, and often in…