activity
20242026
collaborators

6 papers

cs.LG2026

Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks

Binchuan Qi

Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully e…

stat.ML2026

Conjugate Learning Theory: Uncovering the Mechanisms of Trainability and Generalization in Deep Neural Networks

Binchuan Qi

In this work, we propose a notion of practical learnability grounded in finite sample settings, and develop a conjugate learning theoretical framework based on convex conjugate dua…

cs.LG2025

Probability Distribution Learning and Its Application in Deep Learning

Binchuan Qi, Wei Gong, Li Li

Despite its empirical success, deep learning still lacks a comprehensive theoretical understanding of model fitting and generalization. This paper proposes the probability distribu…

cs.LG2025

Extended convexity and smoothness and their applications in deep learning

Binchuan Qi, Wei Gong, Li Li

Classical assumptions like strong convexity and Lipschitz smoothness often fail to capture the nature of deep learning optimization problems, which are typically non-convex and non…

cs.LG2025

Towards Understanding the Optimization Mechanisms in Deep Learning

Binchuan Qi, Wei Gong, Li Li

In this paper, we adopt a probability distribution estimation perspective to explore the optimization mechanisms of supervised classification using deep neural networks. We demonst…

cs.LG2024

Error Bounds of Supervised Classification from Information-Theoretic Perspective

Binchuan Qi

In this paper, we explore bounds on the expected risk when using deep neural networks for supervised classification from an information theoretic perspective. Firstly, we introduce…