Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods
Yotam Wolf, Noam Wies, Dorin Shteyman +3
Language model alignment has become an important component of AI safety, allowing safe interactions between humans and language models, by enhancing desired behaviors and inhibitin…
cs.CL2024
Fundamental Limitations of Alignment in Large Language Models
Yotam Wolf, Noam Wies, Oshri Avnery +2
An important aspect in developing language models that interact with humans is aligning their behavior to be useful and unharmful for their human users. This is usually achieved by…
cs.CL2024
STEER: Assessing the Economic Rationality of Large Language Models
Narun Raman, Taylor Lundy, Samuel Amouyal +3
There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it…