works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CL2026

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexp…

cs.LG2026

On-Policy Delta Distillation

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1

The paper proposes On-Policy Delta Distillation (OPD²), a new on‑policy distillation method that uses a delta signal—the difference between a teacher LLM and its pre‑tuned base mod…

cs.CL2026

Oops, Wait: Discourse Tokens Matter in Reasoning Model

Jaehui Hwang, Byeongho Heo, Sangdoo Yun +1

Recent studies suggest that even data-efficient training with (1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through pos…

cs.IR2026

MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model

Geonmo Gu, Byeongho Heo, Jaemyung Yu +7

Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-targ…

cs.AI2025

What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation

Heejin Do, Jaehui Hwang, Dongyoon Han +2

Evaluating large language models (LLMs) on final-answer correctness is the dominant paradigm. This approach, however, provides a coarse signal for model improvement and overlooks t…