1 paper
Yu Zhao, Yuanbin Qu, Konrad Staniszewski +5
Most language model pre-training frameworks concatenate multiple documents into fixed-length sequences and use causal masking to compute the likelihood of each token given its cont…