Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Are Large Language Models Sensitive to the Motives Behind Communication?
Addison J. Wu, Ryan Liu, Kerem Oktar +2
Human communication is motivated: people speak, write, and create content with a particular communicative intent in mind. As a result, information that large language models (LLMs)…
cs.CL2025
Rational Metareasoning for Large Language Models
C. Nicolò De Sabbata, Theodore R. Sumers, Badr AlKhamissi +2
Being prompted to engage in reasoning has emerged as a core technique for using large language models (LLMs), deploying additional inference-time compute to improve task performanc…
cs.CL2025
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…