Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
Chengtao Jian, Kai Yang, Tianhao Gao +5
Direct Preference Learning has emerged as a dominant offline paradigm for preference optimization. Most of these methods are based on the Bradley-Terry (BT) model for pairwise pref…
cs.AI2025
From Language to Logic: A Bi-Level Framework for Structured Reasoning
Keying Yang, Hao Wang, Kai Yang
Structured reasoning over natural language inputs remains a core challenge in artificial intelligence, as it requires bridging the gap between unstructured linguistic expressions a…