From the 1 of 13 linked papers with an AI index.
10 papers · 1 filter
Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation
Cesare Spinoso-Di Piano, Verna Dankers, Marius Mosbach +1
Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. I…
Value Drifts: Tracing Value Alignment During LLM Post-Training
Mehar Bhatia, Shravan Nayak, Gaurav Kamath +4
The paper studies how large language models acquire and change their alignment with human values during post‑training, analyzing the impact of supervised fine‑tuning and preference…
LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
Matteo Boglioni, Thibault Rousset, Siva Reddy +2
LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged a…
Leveraging Routing Dynamics in Mixture-of-Experts Models for Efficient Language Adaptation
Aditi Khandelwal, Marius Mosbach, Verna Dankers +2
Mixture-of-Experts (MoE) models are widely used to scale language models, yet their expert routing behavior and adaptation in a multilingual setting remain underexplored. In this w…
Forecasting Downstream Performance of LLMs With Proxy Metrics
Arkil Patel, Siva Reddy, Marius Mosbach +1
Progress in language model development is often driven by comparative decisions: which architecture to adopt, which pretraining corpus to use, or which training recipe to apply. Ma…
LLM2Vec-Gen: Generative Embeddings from Large Language Models
Parishad BehnamGhader, Vaibhav Adlakha, Fabian David Schmidt +3
Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output semantics. We propose LLM2Vec-G…