3 papers
cs.AI2026
Emergent Misalignment Is Not Magical
Mingxuan Li, Qirun Dai, Heran Wang +1
Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). EM poses a challenge for A…
cs.LG2025
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
Aniket Vashishtha, Qirun Dai, Hongyuan Mei +3
Counterfactual reasoning, a hallmark of intelligence, consists of three steps: inferring latent variables from observations (abduction), constructing alternatives (interventions),…
cs.CL2025
The Best Instruction-Tuning Data are Those That Fit
Dylan Zhang, Qirun Dai, Hao Peng
High-quality supervised fine-tuning (SFT) data are crucial for eliciting strong capabilities from pretrained large language models (LLMs). Typically, instructions are paired with m…