2 papers
cs.LG2026
Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals
Guopeng Li, Yiyang Duan, Yiru Jiao +1
Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure…
math.OC2024
Zeroth-Order Feedback Optimization in Multi-Agent Systems: Tackling Coupled Constraints
Yingpeng Duan, Yujie Tang
This paper investigates distributed zeroth-order feedback optimization in multi-agent systems with coupled constraints, where each agent operates its local action vector and observ…