2 papers
cs.LG2026
IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
Wanli Zhong, Haibo Feng, Zirui Zhou +2
Deploying Transformer models on edge devices is limited by latency and energy budgets. While INT8 quantization effectively accelerates the primary matrix multiplications, it expose…
cs.LG2026
LLaDA2.1: Speeding Up Text Diffusion via Token Editing
Tiwei Bie, Maosong Cao, Xiang Cao +47
While LLaDA2.0 showcased the scaling potential of 100B-level block-diffusion models and their inherent parallelization, the delicate equilibrium between decoding speed and generati…