Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch
Sen Yang, Yuen-Hei Yeung
An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliabl…
cs.LG2026
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
Sen Yang, Yuen-Hei Yeung
Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this crit…