3 papers
cs.DC2026
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference
Wenchen Han, Gingfung Matthew Yeung, Marco Barletta +3
Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. In disaggregated inference, these…
cs.LG2025
Performance of Zero-Shot Time Series Foundation Models on Cloud Data
William Toner, Thomas L. Lee, Artjom Joosen +2
Time series foundation models (FMs) have emerged as a popular paradigm for zero-shot multi-domain forecasting. FMs are trained on numerous diverse datasets and claim to be effectiv…
cs.LG2025
Lightweight Online Adaption for Time Series Foundation Model Forecasts
Thomas L. Lee, William Toner, Rajkarn Singh +2
Foundation models (FMs) have emerged as a promising approach for time series forecasting. While effective, FMs typically remain fixed during deployment due to the high computationa…