activity
20242026
collaborators

20 papers

cs.LG2026

Transformer Approximations from ReLUs

Jerry Yao-Chieh Hu, Mingcheng Lu, Yi-Chen Lee +1

We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation targets. Importantly, it yiel…

cs.LG2026

Discrete Flow Matching Policy Optimization

Maojiang Su, Po-Chung Hsieh, Weimin Wu +4

We introduce Discrete flow Matching policy Optimization (DoMinO), a unified framework for Reinforcement Learning (RL) fine-tuning Discrete Flow Matching (DFM) models under a broad…

cs.LG2025

On Structured State-Space Duality

Jerry Yao-Chieh Hu, Xiwen Zhang, Ali ElSheikh +2

Structured State-Space Duality (SSD) [Dao & Gu, ICML 2024] is an equivalence between a simple Structured State-Space Model (SSM) and a masked attention mechanism. In particular, a…

cs.LG2025

Universal Approximation with Softmax Attention

Jerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen +2

We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for contin…

cs.LG2025

On Flow Matching KL Divergence

Maojiang Su, Jerry Yao-Chieh Hu, Sophia Pi +1

We derive a deterministic, non-asymptotic upper bound on the Kullback-Leibler (KL) divergence of the flow-matching distribution approximation. In particular, if the flow-matc…

cs.LG2025

A Theoretical Analysis of Discrete Flow Matching Generative Models

Maojiang Su, Mingcheng Lu, Jerry Yao-Chieh Hu +4

We provide a theoretical analysis for end-to-end training Discrete Flow Matching (DFM) generative models. DFM is a promising discrete generative modeling framework that learns the…