activity
20242026
most citedTrustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

1 citations · 1 across the 22 of their papers we have counts for

collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2026

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Prakhar Gupta, Terry Jingchen Zhang, Florent Draye +2

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant…

cs.CL2026

Can Theoretical Physics Research Benefit from Language Agents?

Sirui Lu, Zhijing Jin, Terry Jingchen Zhang +3

Large Language Models (LLMs) are rapidly advancing across diverse domains, yet their application in theoretical physics remains inadequate. While current models show competence in…

cs.CL2026

When Do Language Models Endorse Limitations on Human Rights Principles?

Keenan Samway, Nicole Miu Takagi, Rada Mihalcea +4

As Large Language Models (LLMs) increasingly mediate global information access with the potential to shape public discourse, their alignment with universal human rights principles…

cs.CL2026

Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders

Abir Harrasse, Florent Draye, Punya Syon Pandey +2

Multilingual Large Language Models (LLMs) can process many languages, yet how they internally represent this diversity remains unclear. Do they form shared multilingual representat…

cs.CL2026

On the Emergence and Test-Time Use of Structural Information in Large Language Models

Michelle Chao Chen, Moritz Miller, Bernhard Schölkopf +1

Learning structural information from observational data is central to producing new knowledge outside the training corpus. This holds for mechanistic understanding in scientific di…

cs.CL2025

Improving Large Language Model Safety with Contrastive Representation Learning

Samuel Simko, Mrinmaya Sachan, Bernhard Schölkopf +1

Large Language Models (LLMs) are powerful tools with profound societal impacts, yet their ability to generate responses to diverse and uncontrolled inputs leaves them vulnerable to…