instruction tuning 1large language models 1on-policy distillation 1reasoning 1reinforcement learning 1
From the 1 of 15 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
On-Policy Delta Distillation
Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1
The paper proposes On-Policy Delta Distillation (OPD²), a new on‑policy distillation method that uses a delta signal—the difference between a teacher LLM and its pre‑tuned base mod…
cs.LG2025
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
Song Park, Sanghyuk Chun, Byeongho Heo +1
This paper argues that deep neural networks (DNNs) mostly determine their outputs during the early stages of inference, where biases inherent in the model play a crucial role in sh…
cs.LG2024
Similarity of Neural Architectures using Adversarial Attack Transferability
Jaehui Hwang, Dongyoon Han, Byeongho Heo +3
In recent years, many deep neural architectures have been developed for image classification. Whether they are similar or dissimilar and what factors contribute to their (dis)simil…