3 citations · 4 across the 12 of their papers we have counts for
10 papers · 1 filter
Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space
Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3
Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology. Existing convergence theory, however, exh…
SymPlex: A Structure-Aware Transformer for Symbolic PDE Solving
Yesom Park, Annie C. Lu, Shao-Ching Huang +3
We propose SymPlex, a reinforcement learning framework for discovering analytical symbolic solutions to partial differential equations (PDEs) without access to ground-truth express…
Dynamical Implicit Neural Representations
Yesom Park, Kelvin Kan, Thomas Flynn +4
Implicit Neural Representations (INRs) provide a powerful continuous framework for modeling complex visual and geometric signals, but spectral bias remains a fundamental challenge,…
Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with Prior
Fuqun Han, Stanley Osher, Wuchen Li
In this work, we propose a sparse transformer architecture that incorporates prior information about the underlying data distribution directly into the transformer structure of the…
Stability of Transformers under Layer Normalization
Kelvin Kan, Xingjian Li, Benjamin J. Zhang +4
Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been a…
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3
We study Transformers through the perspective of optimal control theory, using tools from continuous-time formulations to derive actionable insights into training and architecture…