3 papers
cs.LG2026
When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE
Melihcan Erol, Suat Evren, Oktay Ozel +3
InfoNCE is the standard contrastive learning objective, but its softmax form is not only a computational convenience: it also encodes a statistical assumption about how the top-sco…
cs.LG2026
Beyond RLHF: A Unified Theoretical Framework of Alignment
Jihun Yun, Juno Kim, Jongho Park +4
Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However,…
cs.LG2025
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
J. Jon Ryu, Jeongyeol Kwon, Benjamin Koppe +1
We consider off-policy selection and learning in contextual bandits, where the learner aims to select or train a reward-maximizing policy using data collected by a fixed behavior p…