3 papers
cs.DC2026
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference
Wenchen Han, Gingfung Matthew Yeung, Marco Barletta +3
Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. In disaggregated inference, these…
cs.LG2025
Lightweight Online Adaption for Time Series Foundation Model Forecasts
Thomas L. Lee, William Toner, Rajkarn Singh +2
Foundation models (FMs) have emerged as a promising approach for time series forecasting. While effective, FMs typically remain fixed during deployment due to the high computationa…
cs.LG2025
Performance of Zero-Shot Time Series Foundation Models on Cloud Data
William Toner, Thomas L. Lee, Artjom Joosen +2
Time series foundation models (FMs) have emerged as a popular paradigm for zero-shot multi-domain forecasting. FMs are trained on numerous diverse datasets and claim to be effectiv…