2 papers
cs.LG2026
Controllable Value Alignment in Large Language Models through Neuron-Level Editing
Yonghui Yang, Yihui Wang, Junwei Li +6
Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands. However, existing steeri…
cs.LG2026
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
Yonghui Yang, Wenjian Tao, Jilong Liu +6
Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignm…