2 papers
cs.CL2026
Consolidating Rewarded Perturbations for LLM Post-Training
Zheyu Zhang, Shuo Yang, Gjergji Kasneci
Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by RandOpt, relocates this loo…
cs.AI2026
Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language Models
Chenchen Yuan, Zheyu Zhang, Gjergji Kasneci
Large language models often display heterogeneous moral preferences across settings. We study inference-time steering toward a desired ethical framework while preserving general co…