4 papers
Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences
Zhiyu An, Duaa Nakshbandi, Wan Du
Reinforcement learning from human feedback (RLHF) implicitly aggregates heterogeneous human preferences into a single utility function, even though the underlying utilities of the…
DIML: Differentiable Inverse Mechanism Learning from Behaviors of Multi-Agent Learning Trajectories
Zhiyu An, Wan Du
We study inverse mechanism learning: recovering an unknown incentive-generating mechanism from observed strategic interaction traces of self-interested learning agents. Unlike inve…
MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning
Zhiyu An, Wan Du
Large language models are increasingly influencing human moral decisions, yet current approaches focus primarily on evaluating rather than actively steering their moral decisions.…
Disentangling Uncertainties by Learning Compressed Data Representation
Zhiyu An, Zhibo Hou, Wan Du
We study aleatoric and epistemic uncertainty estimation in a learned regressive system dynamics model. Disentangling aleatoric uncertainty (the inherent randomness of the system) f…