134 citations · 458 across the 87 of their papers we have counts for
4 papers · 2 filters
Proto-lm: A Prototypical Network-Based Framework for Built-in Interpretability in Large Language Models
Sean Xie, Soroush Vosoughi, Saeed Hassanpour
Large Language Models (LLMs) have significantly advanced the field of Natural Language Processing (NLP), but their lack of interpretability has been a major concern. Current method…
Training Socially Aligned Language Models on Simulated Social Interactions
Ruibo Liu, Ruixin Yang, Chenyan Jia +5
Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments thr…
Capturing Topic Framing via Masked Language Modeling
Xiaobo Guo, Weicheng Ma, Soroush Vosoughi
Differential framing of issues can lead to divergent world views on important issues. This is especially true in domains where the information presented can reach a large audience,…
Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits
Ruibo Liu, Chenyan Jia, Ge Zhang +3
We present Second Thought, a new learning paradigm that enables language models (LMs) to re-align with human values. By modeling the chain-of-edits between value-unaligned and valu…