5 papers
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
Zhongjie Shi, Wenjing Liao
This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain and -dimensional compact Riemannian manifolds.…
High-Probability Convergence Theory for Distributed Composite Optimization with Sub-Weibull Noises
Zhan Yu, Zhongjie Shi, Deming Yuan
With the rapid development of distributed optimization (DO) theory, the distributed stochastic gradient methods (DSGMs) occupy an important position. Although the theory of differe…
Generic Frameworks for Distributed Functional Optimization and Learning over Time-Varying Networks
Zhan Yu, Zhongjie Shi, Deming Yuan +1
In this paper, we establish a distributed functional optimization (DFO) theory over time-varying networks. The vast majority of existing distributed optimization theories are devel…
Theory of Decentralized Robust Kernel-Based Learning
Zhan Yu, Zhongjie Shi, Ding-Xuan Zhou
We propose a new decentralized robust kernel-based learning algorithm within the framework of reproducing kernel Hilbert spaces (RKHSs) by utilizing a networked system that can be…
Nonlinear functional regression by functional deep neural network with kernel embedding
Zhongjie Shi, Jun Fan, Linhao Song +2
Recently, deep learning has been widely applied in functional data analysis (FDA) with notable empirical success. However, the infinite dimensionality of functional data necessitat…