5 papers
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
Yifan Wang, Jinyi Mu, Mayank Jobanputra +4
Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward model…
A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models
Soyoung Oh, Vera Demberg
To interpret context correctly and retrieve relevant information, large language models must bind entities to their attributes and update these bindings as state changes. We analyz…
Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling
Yifan Wang, Jinyi Mu, Mayank Jobanputra +5
Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values. However, most exist…
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
Yifan Wang, Mayank Jobanputra, Ji-Ung Lee +3
Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes…
Tug-of-war between idioms' figurative and literal interpretations in LLMs
Soyoung Oh, Xinting Huang, Mathis Pink +2
Idioms present a unique challenge for language models due to their non-compositional figurative interpretations, which often strongly diverge from the idiom's literal interpretatio…