8 citations · 30 across the 6 of their papers we have counts for
5 papers · 1 filter
Are Large Language Models Sensitive to the Motives Behind Communication?
Addison J. Wu, Ryan Liu, Kerem Oktar +2
Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs)…
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…
Rational Metareasoning for Large Language Models
C. Nicolò De Sabbata, Theodore R. Sumers, Badr AlKhamissi +2
Being prompted to engage in reasoning has emerged as a core technique for using large language models (LLMs), deploying additional inference-time compute to improve task performanc…
Extending rational models of communication from beliefs to actions
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho +1
Speakers communicate to influence their partner's beliefs and shape their actions. Belief- and action-based objectives have been explored independently in recent computational mode…
Show or Tell? Demonstration is More Robust to Changes in Shared Perception than Explanation
Theodore R. Sumers, Mark K. Ho, Thomas L. Griffiths
Successful teaching entails a complex interaction between a teacher and a learner. The teacher must select and convey information based on what they think the learner perceives and…