activity
20242026
collaborators

14 papers

cs.LG2026

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models

Dake Bu, Wei Huang, Andi Han +5

Discrete diffusion models admit many token orders, yet most systems rely on confidence-based decoding. Confidence is a strong and efficient heuristic, but it can be myopic because…

cs.LG2026

Distributional Biases in Post-Training: A Markovian Analysis of Reasoning Trajectories

Dake Bu, Wei Huang, Andi Han +5

Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiable rewards (RLVR) and test-time scaling (…

cs.LG2026

Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training

Dake Bu, Wei Huang, Andi Han +4

Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a pri…

cs.LG2026

On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD

Tongcheng Zhang, Zhanpeng Zhou, Mingze Wang +4

One crucial factor behind the success of deep learning lies in the implicit bias induced by noise inherent in gradient-based training algorithms. Motivated by empirical observation…

cs.LG2026

Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning

Junsoo Oh, Wei Huang, Taiji Suzuki

Mamba, a recently proposed linear-time sequence model, has attracted significant attention for its computational efficiency and strong empirical performance. However, a rigorous th…

stat.ML2026

Test time training enhances in-context learning of nonlinear functions

Kento Kuwataka, Taiji Suzuki

Test-time training (TTT) enhances model performance by explicitly updating designated parameters prior to each prediction to adapt to the test data. While TTT has demonstrated cons…