1 citations · 1 across the 11 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Fairness Aware Reward Optimization
Ching Lam Choi, Vighnesh Subramaniam, Phillip Isola +2
Demographic skews in human preference data propagate systematic unfairness through reward models into aligned LLMs. We introduce Fairness Aware Reward Optimization (Faro), an in-pr…
cs.LG2024
Characterizing Model Robustness via Natural Input Gradients
Adrián Rodríguez-Muñoz, Tongzhou Wang, Antonio Torralba
Adversarially robust models are locally smooth around each data sample so that small perturbations cannot drastically change model outputs. In modern systems, such smoothness is us…