3 papers
cs.LG2026
Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank
Yiheng Tao, Yihe Zhang, Matthew Dearing +4
Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge that is becoming increasingly acute with t…
cs.LG2025
SMART: A Surrogate Model for Predicting Application Runtime in Dragonfly Systems
Xin Wang, Pietro Lodi Rizzini, Sourav Medya +1
The Dragonfly network, with its high-radix and low-diameter structure, is a leading interconnect in high-performance computing. A major challenge is workload interference on shared…
cs.LG2025
SST: Multi-Scale Hybrid Mamba-Transformer Experts for Time Series Forecasting
Xiongxiao Xu, Canyu Chen, Yueqing Liang +4
Time series forecasting has made significant advances, including with Transformer-based models. The attention mechanism in Transformer effectively captures temporal dependencies by…