9 citations · 18 across the 4 of their papers we have counts for
5 papers · 1 filter
BYOM: Building Your Own Multi-Task Model For Free
Weisen Jiang, Baijiong Lin, Han Shi +3
Recently, various merging methods have been proposed to build a multi-task model from task-specific finetuned models without retraining. However, existing methods suffer from a lar…
Revisiting Over-smoothing in BERT from the Perspective of Graph
Han Shi, Jiahui Gao, Hang Xu +5
Recently over-smoothing phenomenon of Transformer-based models is observed in both vision and language fields. However, no existing work has delved deeper to further investigate th…
SparseBERT: Rethinking the Importance Analysis in Self-attention
Han Shi, Jiahui Gao, Xiaozhe Ren +4
Transformer-based models are popularly used in natural language processing (NLP). Its core component, self-attention, has aroused widespread interest. To understand the self-attent…
Effective Decoding in Graph Auto-Encoder using Triadic Closure
Han Shi, Haozheng Fan, James T. Kwok
The (variational) graph auto-encoder and its variants have been popularly used for representation learning on graph-structured data. While the encoder is often a powerful graph con…
Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONAS
Han Shi, Renjie Pi, Hang Xu +3
Neural Architecture Search (NAS) has shown great potentials in finding better neural network designs. Sample-based NAS is the most reliable approach which aims at exploring the sea…