3 papers
cs.OS2026
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate
Fangyue Liu, Hua Liu, Xinyuan Lyu +8
LLM inference powers latency-critical production services nowadays. The bursty nature of inference traffic results in over-provisioning, which in turn leads to resource underutiliz…
cond-mat.stat-mech2025
Revealing Liquid-Gas Transitions with Finite-Size Scaling in Confined Systems
Chong Zha, Yanshuang Chen, Cheng-Ran Du +2
The application of an external field often renders empirical criteria for identifying liquid-gas phase transitions ambiguous. Here, we demonstrate that the finite-size scaling of t…
cs.CL2025
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought
Tencent Hunyuan Team, Ao Liu, Botong Zhou +248
As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mam…