3 papers
cs.AI2026
Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
Chengtao Jian, Kai Yang, Tianhao Gao +5
Direct Preference Learning has emerged as a dominant offline paradigm for preference optimization. Most of these methods are based on the Bradley-Terry (BT) model for pairwise pref…
cs.AI2025
From Language to Logic: A Bi-Level Framework for Structured Reasoning
Keying Yang, Hao Wang, Kai Yang
Structured reasoning over natural language inputs remains a core challenge in artificial intelligence, as it requires bridging the gap between unstructured linguistic expressions a…
cs.LG2025
Argus: Federated Non-convex Bilevel Learning over 6G Space-Air-Ground Integrated Network
Ya Liu, Kai Yang, Yu Zhu +2
The space-air-ground integrated network (SAGIN) has recently emerged as a core element in the 6G networks. However, traditional centralized and synchronous optimization algorithms…