collaborators

5 papers

cs.AI2026

Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

Yifan Wang, Jinyi Mu, Mayank Jobanputra +4

Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward model…

cs.CL2026

A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models

Soyoung Oh, Vera Demberg

To interpret context correctly and retrieve relevant information, large language models must bind entities to their attributes and update these bindings as state changes. We analyz…

cs.LG2026

Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

Yifan Wang, Jinyi Mu, Mayank Jobanputra +5

Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values. However, most exist…

cs.CL2026

Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?

Yifan Wang, Mayank Jobanputra, Ji-Ung Lee +3

Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes…

cs.CL2026

Tug-of-war between idioms' figurative and literal interpretations in LLMs

Soyoung Oh, Xinting Huang, Mathis Pink +2

Idioms present a unique challenge for language models due to their non-compositional figurative interpretations, which often strongly diverge from the idiom's literal interpretatio…