activity
20242026
collaborators

6 papers

cs.AI2026

Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

Yifan Wang, Jinyi Mu, Mayank Jobanputra +4

Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward model…

cs.LG2026

Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

Yifan Wang, Jinyi Mu, Mayank Jobanputra +5

Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values. However, most exist…

cs.CL2026

Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?

Yifan Wang, Mayank Jobanputra, Ji-Ung Lee +3

Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes…

cs.CL2025

B-cos LM: Efficiently Transforming Pre-trained Language Models for Improved Explainability

Yifan Wang, Sukrut Rao, Ji-Ung Lee +2

Post-hoc explanation methods for black-box models often struggle with faithfulness and human interpretability due to the lack of explainability in current neural architectures. Mea…

cs.HC2025

How AI Responses Shape User Beliefs: The Effects of Information Detail and Confidence on Belief Strength and Stance

Zekun Wu, Mayank Jobanputra, Vera Demberg +2

The growing use of AI-generated responses in everyday tools raises concern about how subtle features such as supporting detail or tone of confidence may shape people's beliefs. To…

cs.LG2025

Can LLMs subtract numbers?

Mayank Jobanputra, Nils Philipp Walter, Maitrey Mehta +7

We present a systematic study of subtraction in large language models (LLMs). While prior benchmarks emphasize addition and multiplication, subtraction has received comparatively l…