activity
20212026
collaborators

8 papers

stat.ML2026

Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity

Zhongjie Shi, Wenjing Liao

This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain and -dimensional compact Riemannian manifolds.…

math.OC2025

Generic Frameworks for Distributed Functional Optimization and Learning over Time-Varying Networks

Zhan Yu, Zhongjie Shi, Deming Yuan +1

In this paper, we establish a distributed functional optimization (DFO) theory over time-varying networks. The vast majority of existing distributed optimization theories are devel…

math.OC2025

High-Probability Convergence Theory for Distributed Composite Optimization with Sub-Weibull Noises

Zhan Yu, Zhongjie Shi, Deming Yuan

With the rapid development of distributed optimization (DO) theory, the distributed stochastic gradient methods (DSGMs) occupy an important position. Although the theory of differe…

cs.LG2025

Theory of Decentralized Robust Kernel-Based Learning

Zhan Yu, Zhongjie Shi, Ding-Xuan Zhou

We propose a new decentralized robust kernel-based learning algorithm within the framework of reproducing kernel Hilbert spaces (RKHSs) by utilizing a networked system that can be…

stat.ML2024

Nonlinear functional regression by functional deep neural network with kernel embedding

Zhongjie Shi, Jun Fan, Linhao Song +2

Recently, deep learning has been widely applied in functional data analysis (FDA) with notable empirical success. However, the infinite dimensionality of functional data necessitat…

stat.ML2023

Learning Theory of Distribution Regression with Neural Networks

Zhongjie Shi, Zhan Yu, Ding-Xuan Zhou

In this paper, we aim at establishing an approximation theory and a learning theory of distribution regression via a fully connected neural network (FNN). In contrast to the classi…