2 citations · 5 across the 3 of their papers we have counts for
4 papers
Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language Model
Mingqi Li, Fei Ding, Dan Zhang +3
Pre-trained multilingual language models play an important role in cross-lingual natural language understanding tasks. However, existing methods did not focus on learning the seman…
Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
Fei Ding, Yin Yang, Hongxin Hu +2
Knowledge distillation (KD) has become an important technique for model compression and knowledge transfer. In this work, we first perform a comprehensive analysis of the knowledge…
Second-order Neural Network Training Using Complex-step Directional Derivative
Siyuan Shen, Tianjia Shao, Kun Zhou +3
While the superior performance of second-order optimization methods such as Newton's method is well known, they are hardly used in practice for deep learning because neither assemb…
CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior
Xiang Zhang, Nishant Vishwamitra, Hongxin Hu +1
We introduce a new deep convolutional neural network, CrescendoNet, by stacking simple building blocks without residual connections. Each Crescendo block contains independent convo…