1 paper · 1 filter
Dong Liu, Yanxuan Yu
As generative models scale to larger inputs across language, vision, and video domains, the cost of token-level computation has become a key bottleneck. While prior work suggests t…