2 papers
cs.LG2026
DOW-KE: Anchor-Free Multi-Layer Knowledge Editing via Direct End-to-End Weight Optimization
Ran Chen, Junbo Zhang, Qianli Zhou +2
Multi-layer locate-then-edit methods for knowledge editing first optimize target residual-stream activations (anchors) at selected layers, then realize them layer by layer as weigh…
cs.CR2025
Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention
Junbo Zhang, Ran Chen, Qianli Zhou +2
Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable app…