3 papers
cs.AI2026
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
Bo Liu, Muxuab Yu, Yu Zhang +2
Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-p…
cs.LG2026
S2O: Early Stopping for Sparse Attention via Online Permutation
Yu Zhang, Songwei Liu, Chenqian Yan +4
Attention scales quadratically with sequence length, fundamentally limiting long-context inference. Existing block-granularity sparsification can reduce latency, but coarse blocks…
cs.CV2025
ARFlow: Autoregressive Flow with Hybrid Linear Attention
Mude Hui, Rui-Jie Zhu, Songlin Yang +5
Flow models are effective at progressively generating realistic images, but they generally struggle to capture long-range dependencies during the generation process as they compres…