activity
20242026
collaborators

6 papers

cs.AI2026

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

Zichao Yu, Chengzhi Yu, Shengze Xu +4

On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repet…

cs.CL2026

Relative Score Policy Optimization for Diffusion Language Models

Zichao Yu, Shengze Xu, Bingqing Jiang +2

Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability requires effective post-training. R…

cs.LG2026

Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow

Hanru Bai, Yuncheng Zhou, Difan Zou

Deep learning paradigms, such as PINNs and neural operators, have significantly advanced the solving of PDEs. However, they often struggle to capture the continuous integral nature…

cs.RO2026

Learning Diffusion Policy from Primitive Skills for Robot Manipulation

Zhihao Gu, Ming Yang, Difan Zou +1

Diffusion policies (DP) have recently shown great promise for generating actions in robotic manipulation. However, existing approaches often rely on global instructions to produce…

cs.LG2025

Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization

Yunzhe Hu, Difan Zou, Dong Xu

Transformer-based models have achieved remarkable success, but their core components, Transformer layers, are largely heuristics-driven and engineered from the bottom up, calling f…

cs.LG2024

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models

Yunzhe Hu, Difan Zou, Dong Xu

Deep neural networks have long been criticized for being black-box. To unveil the inner workings of modern neural architectures, a recent work \cite{yu2024white} proposed an inform…