3 papers
cs.LG2026
Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch
Sen Yang, Yuen-Hei Yeung
An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliabl…
cs.LG2026
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
Sen Yang, Yuen-Hei Yeung
Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this crit…
cs.AI2026
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Sen Yang, Yuen-Hei Yeung
Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failur…