5 papers
Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference
Hui Zang, Pengfei Xia, Hong Liu +5
Mixture-of-Experts (MoE) architectures enable language models to achieve unprecedented scale via sparse activation. However, their inference performance is often limited by data mo…
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
Ziyang Zhang, Xinheng Ding, Jiayi Yuan +4
Deterministic inference is increasingly critical for large language model (LLM) applications such as LLM-as-a-judge evaluation, multi-agent systems, and Reinforcement Learning (RL)…
SciTS: Scientific Time Series Understanding and Generation with LLMs
Wen Wu, Ziyang Zhang, Liwei Liu +12
The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents u…
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
Jun Wang, Yunxiang Yao, Wenwei Kuang +11
Large Language Models drive a wide range of modern AI applications but impose substantial challenges on large-scale serving systems due to intensive computation, strict latency con…
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
Shuo Yan, Ruochen Li, Ziming Luo +11
Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…