Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
Prakhar Gupta, Terry Jingchen Zhang, Florent Draye +2
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant…
cs.CL2026
How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution
Vedant Palit, Florent Draye, Terry Jingchen Zhang +2
Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which…
cs.CL2025
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
Abir Harrasse, Florent Draye, Punya Syon Pandey +2
Multilingual Large Language Models (LLMs) can process many languages, yet how they internally represent this diversity remains unclear. Do they form shared multilingual representat…