1 citations · 1 across the 9 of their papers we have counts for
3 papers · 1 filter
Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space
Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3
Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology. Existing convergence theory, however, exh…
Stability of Transformers under Layer Normalization
Kelvin Kan, Xingjian Li, Benjamin J. Zhang +4
Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been a…
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3
We study Transformers through the perspective of optimal control theory, using tools from continuous-time formulations to derive actionable insights into training and architecture…