3 papers
cs.LG2026
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
Shunchang Liu, Xin Chen, Belen Martin Urcelay +1
Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contr…
cs.HC2026
Beyond Labels: Information-Efficient Human-in-the-Loop Learning using Ranking and Selection Queries
Belén MartÃn-Urcelay, Yoonsang Lee, Matthieu R. Bloch +1
Integrating human expertise into machine learning systems often reduces the role of experts to labeling oracles, a paradigm that limits the amount of information exchanged and fail…
eess.IV2025
MANGO: Learning Disentangled Image Transformation Manifolds with Grouped Operators
Brighton Ancelin, Yenho Chen, Peimeng Guan +4
Learning semantically meaningful image transformations (i.e. rotation, thickness, blur) directly from examples can be a challenging task. Recently, the Manifold Autoencoder (MAE) p…