1 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Zhe Li, Zhangyang Gao, Cheng Tan +2
The pre-training architectures of large language models encompass various types, including autoencoding models, autoregressive models, and encoder-decoder models. We posit that any…