1 paper
Kaicheng Xiao, Haotian Li, Liran Dong +1
While linear attention architectures offer efficient inference, compressing unbounded history into a fixed-size memory inherently limits expressivity and causes information loss. T…