3 papers
cs.AI2026
MiniMax Sparse Attention
Xunhao Lai, Weiqi Xu, Yufeng Yang +14
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointl…
cs.LG2026
MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches
Xin Wang, Chi Ma, Shaobin Chen +14
Generative recommendation (GR) offers superior modeling capabilities but suffers from prohibitive inference costs due to the repeated encoding of long user histories. While cross-r…
physics.chem-ph2025
ByteQC: GPU-Accelerated Quantum Chemistry Package for Large-Scale Systems
Zhen Guo, Zigeng Huang, Qiaorui Chen +7
Applying quantum chemistry algorithms to large-scale systems requires substantial computational resources scaled with the system size and the desired accuracy. To address this, Byt…