Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
Shunchang Liu, Xin Chen, Belen Martin Urcelay +1
Preference learning in large language models relies on reward models as proxies for human judgment. However, these models frequently exhibit preference instability, producing contr…
cs.CV2026
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
Sara Ghazanfari, Francesco Croce, Nicolas Flammarion +3
Recent work has shown that eliciting Large Language Models (LLMs) to generate reasoning traces in natural language before answering the user's request can significantly improve the…