counterfactual reasoning 1incentive compatibility 1language model alignment 1report mediation 1sycophancy 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
Sen Yang, Yuen-Hei Yeung
Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this crit…
cs.AI2026
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Sen Yang, Yuen-Hei Yeung
The paper proposes a method to make language models report truthfully by learning counterfactual report mediators that resist non‑evidential pressure while still updating when genu…