5 papers
Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width
Yanming Lai, Defeng Sun, Yang Wang
In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2…
Beyond the Prompt in Large Language Models: Comprehension, In-Context Learning, and Chain-of-Thought
Yuling Jiao, Yanming Lai, Huazhen Lin +3
Large Language Models (LLMs) have demonstrated remarkable proficiency across diverse tasks, exhibiting emergent properties such as semantic prompt comprehension, In-Context Learnin…
Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with Targets
Yanming Lai, Defeng Sun
The tremendous success of Transformer models in fields such as large language models and computer vision necessitates a rigorous theoretical investigation. To the best of our knowl…
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
Yuling Jiao, Yanming Lai, Yang Wang +1
The Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder…
Approximation Bounds for Transformer Networks with Application to Regression
Yuling Jiao, Yanming Lai, Defeng Sun +2
We explore the approximation capabilities of Transformer networks for Hölder and Sobolev functions, and apply these results to address nonparametric regression estimation with dep…