4 papers · 1 filter
How Language Models Process Negation
Zhejian Zhou, Tianyi Zhou, Robin Jia +1
We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong answers to questions involvi…
Conceptual Steganography
Zhejian Zhou, Jonathan May
Language Models (LMs) emit Chains-of-Thought (CoTs) that drive much of their capability. However, the same sequence that carries useful reasoning can also covertly convey messages:…
Language Models Can Predict Their Own Behavior
Dhananjay Ashok, Jonathan May
The text produced by language models (LMs) can exhibit specific `behaviors,' such as a failure to follow alignment training, that we hope to detect and react to during deployment.…
A Little Human Data Goes A Long Way
Dhananjay Ashok, Jonathan May
Faced with an expensive human annotation process, creators of NLP systems increasingly turn to synthetic data generation. While this method shows promise, the extent to which synth…