2 papers
cs.AI2025
AI Deception: Risks, Dynamics, and Controls
Boyuan Chen, Sitong Fang, Jiaming Ji +56
As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an…
cs.LG2025
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
Liang Lin, Zhihao Xu, Junhao Dong +10
Large language model (LLM) alignment faces a critical dilemma when addressing multiple human preferences: improvements in one dimension frequently come at the expense of others, cr…