3 papers
cs.LG2026
Progressive Approximation in Deep Residual Networks: Theory and Validation
Wei Wang, Xiao-Yong Wei, Qing Li
The Universal Approximation Theorem (UAT) guarantees universal function approximation but does not explain how residual models distribute approximation across layers. We reframe re…
cs.LG2025
TeamFormer: Shallow Parallel Transformers with Progressive Approximation
Wei Wang, Xiao-Yong Wei, Qing Li
The widespread 'deeper is better' philosophy has driven the creation of architectures like ResNet and Transformer, which achieve high performance by stacking numerous layers. Howev…
cs.CL2024
Schrodinger's Memory: Large Language Models
Wei Wang, Qing Li
Memory is the foundation of all human activities; without memory, it would be nearly impossible for people to perform any task in daily life. With the development of Large Language…