20 citations · 50 across the 8 of their papers we have counts for
6 papers · 1 filter
How to Fill the Optimum Set? Population Gradient Descent with Harmless Diversity
Chengyue Gong, Lemeng Wu, Qiang Liu
Although traditional optimization methods focus on finding a single optimal solution, most objective functions in modern machine learning problems, especially those in deep learnin…
Centroid Transformers: Learning to Abstract with Attention
Lemeng Wu, Xingchao Liu, Qiang Liu
Self-attention, as the key block of transformers, is a powerful mechanism for extracting features from the inputs. In essence, what self-attention does is to infer the pairwise rel…
Firefly Neural Architecture Descent: a General Approach for Growing Neural Networks
Lemeng Wu, Bo Liu, Peter Stone +1
We propose firefly neural architecture descent, a general framework for progressively and dynamically growing neural networks to jointly optimize the networks' parameters and archi…
Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
Mao Ye, Lemeng Wu, Qiang Liu
Despite the great success of deep learning, recent works show that large deep neural networks are often highly redundant and can be significantly reduced in size. However, the theo…
Splitting Steepest Descent for Growing Neural Architectures
Qiang Liu, Lemeng Wu, Dilin Wang
We develop a progressive training approach for neural networks which adaptively grows the network structure by splitting existing neurons to multiple off-springs. By leveraging a f…
Energy-Aware Neural Architecture Optimization with Fast Splitting Steepest Descent
Dilin Wang, Meng Li, Lemeng Wu +2
Designing energy-efficient networks is of critical importance for enabling state-of-the-art deep learning in mobile and edge settings where the computation and energy budgets are h…