collaborators

9 papers

cs.LG2026

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models

Dake Bu, Wei Huang, Andi Han +5

Discrete diffusion models admit many token orders, yet most systems rely on confidence-based decoding. Confidence is a strong and efficient heuristic, but it can be myopic because…

cs.LG2026

Distributional Biases in Post-Training: A Markovian Analysis of Reasoning Trajectories

Dake Bu, Wei Huang, Andi Han +5

Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiable rewards (RLVR) and test-time scaling (…

cs.LG2026

Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds

Guoji Fu, Taiji Suzuki, Wee Sun Lee +1

Score-based generative models are trained in high-dimensional ambient spaces, yet many data distributions are supported on low-dimensional nonlinear structures. We prove that, for…

cs.LG2026

Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training

Dake Bu, Wei Huang, Andi Han +4

Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a pri…

stat.ML2025

Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization

Zonghao Chen, Atsushi Nitanda, Arthur Gretton +1

We establish the first global convergence result of neural networks for two stage least squares (2SLS) approach in nonparametric instrumental variable regression (NPIV). This is ac…

stat.ML2025

Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble

Atsushi Nitanda, Anzelle Lee, Damian Tan Xing Kai +2

Mean-field Langevin dynamics (MFLD) is an optimization method derived by taking the mean-field limit of noisy gradient descent for two-layer neural networks in the mean-field regim…