6 papers
Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks
Binchuan Qi
Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully e…
Conjugate Learning Theory: Uncovering the Mechanisms of Trainability and Generalization in Deep Neural Networks
Binchuan Qi
In this work, we propose a notion of practical learnability grounded in finite sample settings, and develop a conjugate learning theoretical framework based on convex conjugate dua…
Probability Distribution Learning and Its Application in Deep Learning
Binchuan Qi, Wei Gong, Li Li
Despite its empirical success, deep learning still lacks a comprehensive theoretical understanding of model fitting and generalization. This paper proposes the probability distribu…
Extended convexity and smoothness and their applications in deep learning
Binchuan Qi, Wei Gong, Li Li
Classical assumptions like strong convexity and Lipschitz smoothness often fail to capture the nature of deep learning optimization problems, which are typically non-convex and non…
Towards Understanding the Optimization Mechanisms in Deep Learning
Binchuan Qi, Wei Gong, Li Li
In this paper, we adopt a probability distribution estimation perspective to explore the optimization mechanisms of supervised classification using deep neural networks. We demonst…
Error Bounds of Supervised Classification from Information-Theoretic Perspective
Binchuan Qi
In this paper, we explore bounds on the expected risk when using deep neural networks for supervised classification from an information theoretic perspective. Firstly, we introduce…