works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators

14 papers

cs.CL2026

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexp…

cs.LG2026

On-Policy Delta Distillation

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1

The paper proposes On-Policy Delta Distillation (OPD²), a new on‑policy distillation method that uses a delta signal—the difference between a teacher LLM and its pre‑tuned base mod…

cs.CL2026

Oops, Wait: Discourse Tokens Matter in Reasoning Model

Jaehui Hwang, Byeongho Heo, Sangdoo Yun +1

Recent studies suggest that even data-efficient training with (1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through pos…

cs.CV2026

RL makes MLLMs see better than SFT

Junha Song, Sangdoo Yun, Dongyoon Han +2

A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its immense parameter scale and remarka…

cs.CV2026

Exploring Conditions for Diffusion models in Robotic Control

Heeseong Shin, Byeongho Heo, Dongyoon Han +2

While pre-trained visual representations have significantly advanced imitation learning, they are often task-agnostic as they remain frozen during policy learning. In this work, we…

cs.IR2026

MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model

Geonmo Gu, Byeongho Heo, Jaemyung Yu +7

Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-targ…