2 papers
cs.LG2026
Scheduling LLM Inference with Uncertainty-Aware Output Length Predictions
Haoyu Zheng, Yongqiang Zhang, Fangcheng Fu +7
To schedule LLM inference, the \textit{shortest job first} (SJF) principle is favorable by prioritizing requests with short output lengths to avoid head-of-line (HOL) blocking. Exi…
cs.CL2025
How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
Qiming Zeng, Xiao Yan, Hao Luo +7
By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) to generate quality answers for user quest…