3 papers
cs.LG2026
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim +1
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillatio…
cs.AI2026
STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure Transformer for Offline Multi-task Multi-agent Reinforcement Learning
Jiwon Jeon, Myungsik Cho, Youngchul Sung
Offline multi-agent reinforcement learning (MARL) with multi-task datasets is challenging due to varying numbers of agents across tasks and the need to generalize to unseen scenari…
cs.MA2026
Generalized Per-Agent Advantage Estimation for Multi-Agent Policy Optimization
Seongmin Kim, Giseung Park, Woojun Kim +3
In this paper, we propose a novel framework for multi-agent reinforcement learning that enhances sample efficiency and coordination through accurate per-agent advantage estimation.…