3 papers
cs.CL2026
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
Haoming Wen, Shi Chen, Qingyu Shi +4
Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised…
cs.LG2026
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
Siyuan Liu, Tinghong Chen, Xinghan Li +2
Data selection during supervised fine-tuning (SFT) can critically change the behavior of large language models (LLMs). Although existing work has studied the effect of selecting da…
cs.CL2026
Differences in Text Generated by Diffusion and Autoregressive Language Models
Zeyang Zhang, Chengwei Liang, Xingyan Chen +4
Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated text remain underexplored. We…