58 citations · 175 across the 29 of their papers we have counts for
6 papers · 1 filter
Revisiting Complete Reasoning Traces for Post-Training
Jaehui Hwang, Sangdoo Yun, Byeongho Heo +1
Large language models (LLMs) are often post-trained on pre-collected reasoning trajectories to improve their reasoning capability. Such trajectories tend to be long due to complex,…
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Seogyeong Jeong, Jaehui Hwang, Dongyoon Han +3
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are expl…
Verification-Aware Training for Speculative Decoding
Geonmo Gu, Byeongho Heo, HeeJae Jun +4
Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, which are verified by the target model in a single forward pass…
On-Policy Delta Distillation for Multilingual Math Reasoning
Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexp…
Oops, Wait: Discourse Tokens Matter in Reasoning Model
Jaehui Hwang, Byeongho Heo, Sangdoo Yun +1
Recent studies suggest that even data-efficient training with (1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through pos…
Token-Supervised Value Models for Enhancing Mathematical Problem-Solving Capabilities of Large Language Models
Jung Hyun Lee, June Yong Yang, Byeongho Heo +4
With the rapid advancement of test-time compute search strategies to improve the mathematical problem-solving capabilities of large language models (LLMs), the need for building ro…