works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CL2026

Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

Cesare Spinoso-Di Piano, Verna Dankers, Marius Mosbach +1

Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. I…

cs.CL2026

Value Drifts: Tracing Value Alignment During LLM Post-Training

Mehar Bhatia, Shravan Nayak, Gaurav Kamath +4

The paper studies how large language models acquire and change their alignment with human values during post‑training, analyzing the impact of supervised fine‑tuning and preference…

cs.CL2026

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

Matteo Boglioni, Thibault Rousset, Siva Reddy +2

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged a…

cs.CV2026

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

Benno Krojer, Shravan Nayak, Oscar Mañas +4

Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM.…

cs.LG2026

Operationalising the Superficial Alignment Hypothesis via Task Complexity

Tomás Vergara-Browne, Darshan Patil, Ivan Titov +3

The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that post-training merely surfaces this knowledg…

cs.CL2026

Leveraging Routing Dynamics in Mixture-of-Experts Models for Efficient Language Adaptation

Aditi Khandelwal, Marius Mosbach, Verna Dankers +2

Mixture-of-Experts (MoE) models are widely used to scale language models, yet their expert routing behavior and adaptation in a multilingual setting remain underexplored. In this w…