4 papers
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
Zhou Ziheng, Jiakun Ding, Zhaowei Zhang +6
Principle-based alignment often lacks context sensitivity and completeness. Grounded in Theory of Mind, we propose role conditioning as a compact alternative: social roles (e.g., m…
BACH-V: Bridging Abstract and Concrete Human-Values in Large Language Models
Junyu Zhang, Yipeng Kang, Jiong Guo +2
Do large language models (LLMs) genuinely understand abstract concepts, or merely manipulate them as statistical patterns? We introduce an abstraction-grounding framework that deco…
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
Yipeng Kang, Junqi Wang, Yexin Li +8
As large language models (LLMs) become increasingly integrated into critical applications, aligning their behavior with human values presents significant challenges. Current method…
Score Neural Operator: A Generative Model for Learning and Generalizing Across Multiple Probability Distributions
Xinyu Liao, Aoyang Qin, Jacob Seidman +3
Most existing generative models are limited to learning a single probability distribution from the training data and cannot generalize to novel distributions for unseen data. An ar…