5 papers
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim +1
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillatio…
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Jeonghye Kim, Xufang Luo, Minbeom Kim +5
Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in mathematical reasoning, we…
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Jeonghye Kim, Jiwon Jeon, Dongsheng Li +1
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model…
STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure Transformer for Offline Multi-task Multi-agent Reinforcement Learning
Jiwon Jeon, Myungsik Cho, Youngchul Sung
Offline multi-agent reinforcement learning (MARL) with multi-task datasets is challenging due to varying numbers of agents across tasks and the need to generalize to unseen scenari…
Generalized Per-Agent Advantage Estimation for Multi-Agent Policy Optimization
Seongmin Kim, Giseung Park, Woojun Kim +3
In this paper, we propose a novel framework for multi-agent reinforcement learning that enhances sample efficiency and coordination through accurate per-agent advantage estimation.…