collaborators

28 papers

cs.CL2026

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

Hadas Orgad, Boyi Wei, Kaden Zheng +4

Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate…

cs.LG20261 cited

Two Stages of Folding: Convergent Mechanisms in AI Protein Folding Trunks

Kevin Lu, Jannik Brinkmann, Stefan Huber +4

How do protein structure prediction models fold proteins? We investigate this question through causal interventions on the folding trunks of ESMFold, OpenFold, and Boltz-1. Across…

cs.CL2026

Reasoning Models Know What's Important, and Encode It in Their Activations

Yaniv Nikankin, Martin Tutek, Tomer Ashuach +2

Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the fin…

cs.LG2026

Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models

Gal Pomerants, Yaniv Nikankin, Anja Reusch +3

Protein sequences are abundant in repeating segments, both as exact copies and as approximate segments with mutations. These repeats are important for protein structure and functio…

cs.CL2026

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…

cs.CL2026

Old Habits Die Hard: How Conversational History Geometrically Traps LLMs

Adi Simhi, Fazl Barez, Martin Tutek +2

How does the conversational past of large language models (LLMs) influence their future performance? Recent work suggests that LLMs are affected by their conversational history in…