Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
LLaDA2.1: Speeding Up Text Diffusion via Token Editing
Tiwei Bie, Maosong Cao, Xiang Cao +47
While LLaDA2.0 showcased the scaling potential of 100B-level block-diffusion models and their inherent parallelization, the delicate equilibrium between decoding speed and generati…
cs.LG2026
Towards Compressive and Scalable Recurrent Memory
Yunchong Song, Jushi Kai, Liming Lu +2
Transformers face a quadratic bottleneck in attention when scaling to long contexts. Recent approaches introduce recurrent memory to extend context beyond the current window, yet t…
cs.LG2025
Three-dimensional attention Transformer for state evaluation in real-time strategy games
Yanqing Ye, Weilong Yang, Kai Qiu +1
Situation assessment in Real-Time Strategy (RTS) games is crucial for understanding decision-making in complex adversarial environments. However, existing methods remain limited in…