Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
Shunchang Liu, Xin Chen, Belen Martin Urcelay +1
Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contr…
cs.LG2025
Selective Induction Heads: How Transformers Select Causal Structures In Context
Francesco D'Angelo, Francesco Croce, Nicolas Flammarion
Transformers have exhibited exceptional capabilities in sequence modeling tasks, leveraging self-attention and in-context learning. Critical to this success are induction heads, at…