activity
20162022
most citedAdaptive Multi-Teacher Multi-level Knowledge Distillation

222 citations · 338 across the 23 of their papers we have counts for

collaborators

31 papers

cs.LG20221 cited

DEMAND: Deep Matrix Approximately Nonlinear Decomposition to Identify Meta, Canonical, and Sub-Spatial Pattern of functional Magnetic Resonance Imaging in the Human Brain

Wei Zhang, Yu Bao

Deep Neural Networks (DNNs) have already become a crucial computational approach to revealing the spatial patterns in the human brain; however, there are three major shortcomings i…

cs.LG2022

SADAM: Stochastic Adam, A Stochastic Operator for First-Order Gradient-based Optimizer

Wei Zhang, Yu Bao

In this work, to efficiently help escape the stationary and saddle points, we propose, analyze, and generalize a stochastic strategy performed as an operator for a first-order grad…

cs.CL2022

Detecting Textual Adversarial Examples Based on Distributional Characteristics of Data Representations

Na Liu, Mark Dras, Wei Emma Zhang

Although deep neural networks have achieved state-of-the-art performance in various machine learning tasks, adversarial examples, constructed by adding small non-random perturbatio…

eess.SP2022

Beam Training and Alignment for RIS-Assisted Millimeter Wave Systems:State of the Art and Beyond

Peilan Wang, Jun Fang, Weizheng Zhang +3

Reconfigurable intelligent surface (RIS) has recently emerged as a promising paradigm for future cellular networks. Specifically, due to its capability in reshaping the propagation…

cs.CV2021222 cited

Adaptive Multi-Teacher Multi-level Knowledge Distillation

Yuang Liu, Wei Zhang, Jun Wang

Knowledge distillation~(KD) is an effective learning paradigm for improving the performance of lightweight student networks by utilizing additional supervision knowledge distilled…

cs.LG202012 cited

Adam: A Stochastic Method with Adaptive Variance Reduction

Mingrui Liu, Wei Zhang, Francesco Orabona +1

Adam is a widely used stochastic optimization method for deep learning applications. While practitioners prefer Adam because it requires less parameter tuning, its use is problemat…