9 papers
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models
Dake Bu, Wei Huang, Andi Han +5
Discrete diffusion models admit many token orders, yet most systems rely on confidence-based decoding. Confidence is a strong and efficient heuristic, but it can be myopic because…
Distributional Biases in Post-Training: A Markovian Analysis of Reasoning Trajectories
Dake Bu, Wei Huang, Andi Han +5
Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiable rewards (RLVR) and test-time scaling (…
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
Guoji Fu, Taiji Suzuki, Wee Sun Lee +1
Score-based generative models are trained in high-dimensional ambient spaces, yet many data distributions are supported on low-dimensional nonlinear structures. We prove that, for…
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
Dake Bu, Wei Huang, Andi Han +4
Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a pri…
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
Zonghao Chen, Atsushi Nitanda, Arthur Gretton +1
We establish the first global convergence result of neural networks for two stage least squares (2SLS) approach in nonparametric instrumental variable regression (NPIV). This is ac…
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
Atsushi Nitanda, Anzelle Lee, Damian Tan Xing Kai +2
Mean-field Langevin dynamics (MFLD) is an optimization method derived by taking the mean-field limit of noisy gradient descent for two-layer neural networks in the mean-field regim…