Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
Woohyeon Byeon, Jiwon Jeon, Jeonghye Kim +1
We study multi-domain LLM training in which two models, each stronger in a different domain, co-evolve by tutoring each other through on-policy feedback. Unlike one-way distillatio…
cs.LG2026
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
Jeonghye Kim, Jiwon Jeon, Dongsheng Li +1
Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student without it, both from the same model…