Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
TeamFormer: Shallow Parallel Transformers with Progressive Approximation
Wei Wang, Xiao-Yong Wei, Qing Li
The widespread 'deeper is better' philosophy has driven the creation of architectures like ResNet and Transformer, which achieve high performance by stacking numerous layers. Howev…
cs.LG2024
Dynamic Universal Approximation Theory: Foundations for Parallelism in Neural Networks
Wei Wang, Qing Li
Neural networks are increasingly evolving towards training large models with big data, a method that has demonstrated superior performance across many tasks. However, this approach…