1 citations · 1 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding
Mingxi Zou, Jiaxiang Chen, Junfan Li +4
Preference-based alignment methods (e.g., RLHF, DPO) typically optimize a single scalar objective, implicitly averaging over heterogeneous human preferences. In practice, systemati…
cs.LG2024
Online Optimization for Learning to Communicate over Time-Correlated Channels
Zheshun Wu, Junfan Li, Zenglin Xu +2
Machine learning techniques have garnered great interest in designing communication systems owing to their capacity in tackling with channel uncertainty. To provide theoretical gua…