32 citations · 34 across the 27 of their papers we have counts for
3 papers · 1 filter
Learning-to-Defer in Non-Stationary Time Series via Switching State-Space Models
Yannis Montreuil, Letian Yu, Axel Carlier +2
Learning-to-defer (L2D) routes each decision to a system's own predictor or to an external expert. Streaming time-series settings break the offline-L2D assumptions: the data are no…
Why Ask One When You Can Ask ? Learning-to-Defer to the Top- Experts
Yannis Montreuil, Axel Carlier, Lai Xing Ng +1
Existing Learning-to-Defer (L2D) frameworks are limited to single-expert deferral, forcing each query to rely on only one expert and preventing the use of collective expertise. We…
FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
Flint Xiaofeng Fan, Cheston Tan, Yew-Soon Ong +2
In the era of increasing privacy concerns and demand for personalized experiences, traditional Reinforcement Learning with Human Feedback (RLHF) frameworks face significant challen…