17 papers
On-Policy Delta Distillation for Multilingual Math Reasoning
Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexp…
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
Seunghyun Lee, Dongyoon Han, Sangdoo Yun
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal…
On-Policy Delta Distillation
Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1
The paper proposes On-Policy Delta Distillation (OPD²), a new on‑policy distillation method that uses a delta signal—the difference between a teacher LLM and its pre‑tuned base mod…
Oops, Wait: Discourse Tokens Matter in Reasoning Model
Jaehui Hwang, Byeongho Heo, Sangdoo Yun +1
Recent studies suggest that even data-efficient training with (1K) reasoning trajectories can induce non-trivial reasoning capabilities in large language models through pos…
Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time
Jeongeun Park, Juhan Park, Taekyung Kim +3
Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in bot…
WorldKV: Efficient World Memory with World Retrieval and Compression
Jung Yi, Minjae Kim, Paul Hyunbin Cho +3
Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpo…