collaborators

11 papers

cs.LG2025

A Theoretical Analysis of Discrete Flow Matching Generative Models

Maojiang Su, Mingcheng Lu, Jerry Yao-Chieh Hu +4

We provide a theoretical analysis for end-to-end training Discrete Flow Matching (DFM) generative models. DFM is a promising discrete generative modeling framework that learns the…

cs.LG2025

Are Hallucinations Bad Estimations?

Hude Liu, Jerry Yao-Chieh Hu, Jennifer Yuntong Zhang +2

We formalize hallucinations in generative models as failures to link an estimate to any plausible cause. Under this interpretation, we show that even loss-minimizing optimal estima…

cs.LG2025

Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models

Jerry Yao-Chieh Hu, Maojiang Su, En-Jui Kuo +2

We study the computational limits of Low-Rank Adaptation (LoRA) for finetuning transformer-based models using fine-grained complexity theory. Our key observation is that the existe…

cs.LG2025

Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency

Jerry Yao-Chieh Hu, Wei-Po Wang, Ammar Gilani +3

We investigate the statistical and computational limits of prompt tuning for transformer-based foundation models. Our key contributions are prompt tuning on \emph{single-head} tran…

cs.LG2025

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions

Jerry Yao-Chieh Hu, Xiwen Zhang, Maojiang Su +2

We study the computational limits of learning -bit Boolean functions (specifically, , , and their noisy variants), using a minimalist single-head soft…

cs.LG2025

Attention Mechanism, Max-Affine Partition, and Universal Approximation

Hude Liu, Jerry Yao-Chieh Hu, Zhao Song +1

We establish the universal approximation capability of single-layer, single-head self- and cross-attention mechanisms with minimal attached structures. Our key insight is to interp…