3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Exploring the Benefit of Activation Sparsity in Pre-training
Zhengyan Zhang, Chaojun Xiao, Qiujieli Qin +7
Pre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activatio…
cs.IR2023★ 3 cited
Thoroughly Modeling Multi-domain Pre-trained Recommendation as Language
Zekai Qu, Ruobing Xie, Chaojun Xiao +5
With the thriving of pre-trained language model (PLM) widely verified in various of NLP tasks, pioneer efforts attempt to explore the possible cooperation of the general textual in…