2 papers
cs.AI2026
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
Valen Tagliabue, Leonard Dung
We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with preferences expressed through behav…
cs.CR2024
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan +7
Large Language Models (LLMs) are deployed in interactive contexts with direct user engagement, such as chatbots and writing assistants. These deployments are vulnerable to prompt i…