activity
20242026
most citedNumerical Analysis on Neural Network Projected Schemes for Approximating One Dimensional Wasserstein Gradient Flows

3 citations · 4 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2026

Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space

Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3

Discrete diffusion has become a leading framework for generative modeling in various applications including language, vision, and biology. Existing convergence theory, however, exh…

cs.LG2026

SymPlex: A Structure-Aware Transformer for Symbolic PDE Solving

Yesom Park, Annie C. Lu, Shao-Ching Huang +3

We propose SymPlex, a reinforcement learning framework for discovering analytical symbolic solutions to partial differential equations (PDEs) without access to ground-truth express…

cs.LG2025

Dynamical Implicit Neural Representations

Yesom Park, Kelvin Kan, Thomas Flynn +4

Implicit Neural Representations (INRs) provide a powerful continuous framework for modeling complex visual and geometric signals, but spectral bias remains a fundamental challenge,…

cs.LG2025

Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with Prior

Fuqun Han, Stanley Osher, Wuchen Li

In this work, we propose a sparse transformer architecture that incorporates prior information about the underlying data distribution directly into the transformer structure of the…

cs.LG2025

Stability of Transformers under Layer Normalization

Kelvin Kan, Xingjian Li, Benjamin J. Zhang +4

Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been a…

cs.LG2025

Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency

Kelvin Kan, Xingjian Li, Benjamin J. Zhang +3

We study Transformers through the perspective of optimal control theory, using tools from continuous-time formulations to derive actionable insights into training and architecture…