4 papers
TeamFormer: Shallow Parallel Transformers with Progressive Approximation
Wei Wang, Xiao-Yong Wei, Qing Li
The widespread 'deeper is better' philosophy has driven the creation of architectures like ResNet and Transformer, which achieve high performance by stacking numerous layers. Howev…
Dynamic Universal Approximation Theory: The Basic Theory for Transformer-based Large Language Models
Wei Wang, Qing Li
Language models have emerged as a critical area of focus in artificial intelligence, particularly with the introduction of groundbreaking innovations like ChatGPT. Large-scale Tran…
Dynamic Universal Approximation Theory: Foundations for Parallelism in Neural Networks
Wei Wang, Qing Li
Neural networks are increasingly evolving towards training large models with big data, a method that has demonstrated superior performance across many tasks. However, this approach…
Dynamic Universal Approximation Theory: The Basic Theory for Deep Learning-Based Computer Vision Models
Wei Wang, Qing Li
Computer vision (CV) is one of the most crucial fields in artificial intelligence. In recent years, a variety of deep learning models based on convolutional neural networks (CNNs)…