3 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Bo Liu, Rui Wang, Lemeng Wu +3
Modern large language models are built on sequence modeling via next-token prediction. While the Transformer remains the dominant architecture for sequence modeling, its quadratic…