3 papers
cs.LG2026
Meta-Learning Objectives for Preference Optimization
Carlo Alfano, Silvia Sapora, Jakob Nicolaus Foerster +2
Evaluating preference optimization (PO) algorithms on LLM alignment is a challenging task that presents prohibitive costs, noise, and several variables like model size and hyper-pa…
stat.ML2026
Learning mirror maps in policy mirror descent
Carlo Alfano, Sebastian Towers, Silvia Sapora +2
Policy Mirror Descent (PMD) is a popular framework in reinforcement learning, serving as a unifying perspective that encompasses numerous algorithms. These algorithms are derived t…
cs.CL2025
Multilingual Self-Taught Faithfulness Evaluators
Carlo Alfano, Aymen Al Marjani, Zeno Jonke +3
The growing use of large language models (LLMs) has increased the need for automatic evaluation systems, particularly to address the challenge of information hallucination. Althoug…