26 citations · 31 across the 6 of their papers we have counts for
5 papers · 1 filter
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
Dake Bu, Wei Huang, Andi Han +4
Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a pri…
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
Dake Bu, Wei Huang, Andi Han +4
In-context learning (ICL) has garnered significant attention for its ability to grasp functions/tasks from demonstrations. Recent studies suggest the presence of a latent task/func…
Direct Distributional Optimization for Provable Alignment of Diffusion Models
Ryotaro Kawata, Kazusato Oko, Atsushi Nitanda +1
We introduce a novel alignment method for diffusion models from distribution optimization perspectives while providing rigorous convergence guarantees. We first formulate the probl…
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
Dake Bu, Wei Huang, Andi Han +4
Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection b…
Convergence of mean-field Langevin dynamics: Time and space discretization, stochastic gradient, and variance reduction
Taiji Suzuki, Denny Wu, Atsushi Nitanda
The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the op…