2 papers
stat.ML2025
Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
Shai Zucker, Xiong Wang, Fei Lu +1
We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation fun…
math.ST2025
Minimax rates for learning kernels in operators
Sichong Zhang, Xiong Wang, Fei Lu
Learning kernels in operators from data lies at the intersection of inverse problems and statistical learning, providing a powerful framework for capturing non-local dependencies i…