2 papers
cs.MA2026
ArcticSwarm: Deferring Early Consensus in Long-Horizon Multi-Agent Research
Soyoung Yoon, Boyi Liu, Yite Wang +6
Multi-agent systems have shown strong performance in domains with reliable verifiers such as coding, where multi-parallel candidate generation selected by a verifier is effective.…
cs.LG2023
Posterior Sampling with Delayed Feedback for Reinforcement Learning with Linear Function Approximation
Nikki Lijing Kuang, Ming Yin, Mengdi Wang +2
Recent studies in reinforcement learning (RL) have made significant progress by leveraging function approximation to alleviate the sample complexity hurdle for better performance.…