2 papers
cs.LG2026
Fairness Aware Reward Optimization
Ching Lam Choi, Vighnesh Subramaniam, Phillip Isola +2
Demographic skews in human preference data propagate systematic unfairness through reward models into aligned LLMs. We introduce Fairness Aware Reward Optimization (Faro), an in-pr…
cs.LG2025
Unlearning-based Neural Interpretations
Ching Lam Choi, Alexandre Duplessis, Serge Belongie
Gradient-based interpretations often require an anchor point of comparison to avoid saturation in computing feature importance. We show that current baselines defined using static…