activity
20242026
collaborators

5 papers

stat.ML2026

Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity

Zhongjie Shi, Wenjing Liao

This paper investigates the learning theory of Transformer networks for regression tasks on the compact Euclidean domain and -dimensional compact Riemannian manifolds.…

math.OC2026

High-Probability Convergence Theory for Distributed Composite Optimization with Sub-Weibull Noises

Zhan Yu, Zhongjie Shi, Deming Yuan

With the rapid development of distributed optimization (DO) theory, the distributed stochastic gradient methods (DSGMs) occupy an important position. Although the theory of differe…

math.OC2025

Generic Frameworks for Distributed Functional Optimization and Learning over Time-Varying Networks

Zhan Yu, Zhongjie Shi, Deming Yuan +1

In this paper, we establish a distributed functional optimization (DFO) theory over time-varying networks. The vast majority of existing distributed optimization theories are devel…

cs.LG2025

Theory of Decentralized Robust Kernel-Based Learning

Zhan Yu, Zhongjie Shi, Ding-Xuan Zhou

We propose a new decentralized robust kernel-based learning algorithm within the framework of reproducing kernel Hilbert spaces (RKHSs) by utilizing a networked system that can be…

stat.ML2025

Nonlinear functional regression by functional deep neural network with kernel embedding

Zhongjie Shi, Jun Fan, Linhao Song +2

Recently, deep learning has been widely applied in functional data analysis (FDA) with notable empirical success. However, the infinite dimensionality of functional data necessitat…