6 papers
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
Zichao Yu, Chengzhi Yu, Shengze Xu +4
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repet…
Relative Score Policy Optimization for Diffusion Language Models
Zichao Yu, Shengze Xu, Bingqing Jiang +2
Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability requires effective post-training. R…
Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow
Hanru Bai, Yuncheng Zhou, Difan Zou
Deep learning paradigms, such as PINNs and neural operators, have significantly advanced the solving of PDEs. However, they often struggle to capture the continuous integral nature…
Learning Diffusion Policy from Primitive Skills for Robot Manipulation
Zhihao Gu, Ming Yang, Difan Zou +1
Diffusion policies (DP) have recently shown great promise for generating actions in robotic manipulation. However, existing approaches often rely on global instructions to produce…
Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
Yunzhe Hu, Difan Zou, Dong Xu
Transformer-based models have achieved remarkable success, but their core components, Transformer layers, are largely heuristics-driven and engineered from the bottom up, calling f…
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
Yunzhe Hu, Difan Zou, Dong Xu
Deep neural networks have long been criticized for being black-box. To unveil the inner workings of modern neural architectures, a recent work \cite{yu2024white} proposed an inform…