Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Safe Evolution with Circuit Anchors
Yan Liu, Jie Fu, Tsung-Yi Ho
In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature'…
cs.CL2024
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
Shitong Duan, Xiaoyuan Yi, Peng Zhang +5
Large language models (LLMs) have revolutionized the role of AI, yet pose potential social risks. To steer LLMs towards human preference, alignment technologies have been introduce…
cs.CL2024
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
Yan Liu, Xiaoyuan Yi, Xiaokang Chen +6
The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models…