4 papers
OUTLETS: Output-Length Prediction from Speculative Decoding Backbones
Weihuang Wen, Yingying Liu, Yichuan Liu +5
The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-lengt…
CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean
Wentao Long, Yunfei Zhang, Chenyi Li +3
Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchmarks mainly focus on Olympia…
A Learning Method with Gap-Aware Generation for Heterogeneous DAG Scheduling
Ruisong Zhou, Haijun Zou, Li Zhou +2
Efficient scheduling of directed acyclic graphs (DAGs) is a core problem in large-scale data-intensive computing systems, where query plans, data-processing workloads, and computat…
Synera: Synergistic LLM Serving across Device and Cloud at Scale
Genglin Wang, Liekang Zeng, Bufang Yang +6
Large Language Models (LLMs) are becoming key components in various mobile operating systems, driving smart applications like interactive chatbots and personal assistants. While br…