machine learning

Minimum Block Width for Universal Approximation by Residual Neural Networks with Inner Width One

arXiv:2607.04597

summary

The paper analyzes the minimal block width required for residual neural networks with inner width one to achieve universal function approximation, proving it equals the larger of input or output dimensions and providing bounds for both L^p and uniform approximation.

Abstract

In this paper, we study the universal approximation property of residual neural networks. For input and output dimensions and , and LeakyReLU, ReLU, ReLU-like activation functions, the upper and lower bounds of the minimum block width are established. To achieve approximation on any compact set, we show that the exact minimum block width is when each residual branch has inner width 1. Furthermore, we show that residual neural networks with block width can achieve uniform approximation on any compact set under the constraint that each residual branch has inner width 1. Besides, for any activation function family, we prove that there exist functions that cannot be approximated by residual neural networks with block width less than , both in the sense and the uniform sense, regardless of inner width. Consequently, for LeakyReLU, ReLU, ReLU-like activation functions and , the exact minimum block width for uniform approximation is when each residual branch has inner width 1.

Topics & keywords

#universal approximation#residual networks#network width#function approximation#deep learning theoryresidual neural networkblock widthinner widthLeakyReLUReLUL^p approximationuniform approximation