3 papers
cs.CL2026
HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion
Junxiang Qiu, Zhengsu Chen, Xinting Hu +6
Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and…
cs.CL2026
Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models
Junxiang Qiu, Shuo Wang, Zhengsu Chen +4
Attention serves as the fundamental mechanism for long-context modeling in large language models (LLMs), yet dense attention becomes structurally prohibitive for long sequences due…
cs.CV2025
Accelerating Controllable Generation via Hybrid-grained Cache
Lin Liu, Huixia Ben, Shuo Wang +4
Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation…