2 papers
cs.LG2026
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
Hongli Shen, Shaopeng Fu, Qinbo Zhang +2
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direc…
cs.CL2026
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
Weixu Zhang, Fanghua Ye, Qiang Gao +7
Large language models (LLMs) often produce content that contradicts or overlooks information provided in the input context, a phenomenon known as faithfulness hallucination. In thi…