6 papers
Sparse Layers are Critical to Scaling Looped Language Models
Ryan Lee, Jacob Biloki, Edward J. Hu +1
Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do…
How Language Models Process Negation
Zhejian Zhou, Tianyi Zhou, Robin Jia +1
We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong answers to questions involvi…
Conceptual Steganography
Zhejian Zhou, Jonathan May
Language Models (LMs) emit Chains-of-Thought (CoTs) that drive much of their capability. However, the same sequence that carries useful reasoning can also covertly convey messages:…
Language Models Can Predict Their Own Behavior
Dhananjay Ashok, Jonathan May
The text produced by language models (LMs) can exhibit specific `behaviors,' such as a failure to follow alignment training, that we hope to detect and react to during deployment.…
Can VLMs Recall Factual Associations From Visual References?
Dhananjay Ashok, Ashutosh Chaubey, Hirona J. Arai +2
Through a controlled study, we identify a systematic deficiency in the multimodal grounding of Vision Language Models (VLMs). While VLMs can recall factual associations when provid…
A Little Human Data Goes A Long Way
Dhananjay Ashok, Jonathan May
Faced with an expensive human annotation process, creators of NLP systems increasingly turn to synthetic data generation. While this method shows promise, the extent to which synth…