4 citations · 4 across the 1 of their papers we have counts for
2 papers
cs.CL2021
Training Multilingual Pre-trained Language Model with Byte-level Subwords
Junqiu Wei, Qun Liu, Yinpeng Guo +1
The pre-trained language models have achieved great successes in various natural language understanding (NLU) tasks due to its capacity to capture the deep contextualized informati…
cs.CL2020★ 4 cited
TensorCoder: Dimension-Wise Attention via Tensor Representation for Natural Language Modeling
Shuai Zhang, Peng Zhang, Xindian Ma +3
Transformer has been widely-used in many Natural Language Processing (NLP) tasks and the scaled dot-product attention between tokens is a core module of Transformer. This attention…