2 papers
cs.LG2026
Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank
Yiheng Tao, Yihe Zhang, Matthew Dearing +4
Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with t…
cs.LG2025
SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
Xin Wang, Pietro Lodi Rizzini, Sourav Medya +1
The Dragonfly network, with its high-radix and low-diameter structure, is a leading interconnect in high-performance computing. A major challenge is workload interference on shared…