collaborators

17 papers

cs.CL2026

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexp…

cs.LG2026

Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs

Seunghyun Lee, Dongyoon Han, Sangdoo Yun

Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal…

cs.LG2026

On-Policy Delta Distillation

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1

The paper proposes On-Policy Delta Distillation (OPD²), a new on‑policy distillation method that uses a delta signal—the difference between a teacher LLM and its pre‑tuned base mod…

cs.CL2026

Oops, Wait: Discourse Tokens Matter in Reasoning Model

Jaehui Hwang, Byeongho Heo, Sangdoo Yun +1

Recent studies suggest that even data-efficient training with (1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through pos…

cs.RO2026

Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

Jeongeun Park, Juhan Park, Taekyung Kim +3

Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in bot…

cs.CV2026

WorldKV: Efficient World Memory with World Retrieval and Compression

Jung Yi, Minjae Kim, Paul Hyunbin Cho +3

Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpo…