instruction tuning 1large language models 1on-policy distillation 1reasoning 1reinforcement learning 1
From the 1 of 15 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
On-Policy Delta Distillation for Multilingual Math Reasoning
Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexp…
cs.CL2026
Oops, Wait: Discourse Tokens Matter in Reasoning Model
Jaehui Hwang, Byeongho Heo, Sangdoo Yun +1
Recent studies suggest that even data-efficient training with (1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through pos…
cs.CL2025
Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
Jung Hyun Lee, June Yong Yang, Byeongho Heo +4
With the rapid advancement of test-time compute search strategies to improve the mathematical problem-solving capabilities of large language models (LLMs), the need for building ro…