3 papers
cs.LG2026
Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning
Xinrui Chen, Jianhao Zhang, Ou Wu +1
Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods use fixed safety examples, gl…
cs.AI2026
Revisiting Ripple Effects in Knowledge Editing through Pressure-Aware Joint Neighborhood Optimization
Haoben Huang, Shuxin Liu, Ou Wu +1
Single-edit updates in large language models can trigger ripple effects across local knowledge neighborhoods: desirable propagation to related facts and unintended perturbation of…
cs.CL2026
MetaKE: Meta-Learning for Knowledge Editing Toward a Better Accuracy-Editability Trade-off
Shuxin Liu, Di Gao, Ou Wu
Existing locate-then-edit Knowledge Editing (KE) methods typically decompose editing into two stages: upstream target representation optimization and downstream constrained paramet…