28 papers
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Hadas Orgad, Boyi Wei, Kaden Zheng +4
Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate…
Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks
Kevin Lu, Jannik Brinkmann, Stefan Huber +4
How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across…
Reasoning Models Know What's Important, and Encode It in Their Activations
Yaniv Nikankin, Martin Tutek, Tomer Ashuach +2
Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the fin…
Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models
Gal Pomerants, Yaniv Nikankin, Anja Reusch +3
Protein sequences are abundant in repeating segments, both as exact copies and as approximate segments with mutations. These repeats are important for protein structure and functio…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
Adi Simhi, Fazl Barez, Martin Tutek +2
How does the conversational past of large language models (LLMs) influence their future performance? Recent work suggests that LLMs are affected by their conversational history in…