8 citations · 8 across the 3 of their papers we have counts for
1 paper · 1 filter
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang +4
Computation in a typical Transformer-based large language model (LLM) can be characterized by batch size, hidden dimension, number of layers, and sequence length. Until now, system…