3 papers
cs.DC2026
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
Chen Zhang, Yan Ding, Haotian Wang +3
During the deployment of Large Language Models (LLMs), the autoregressive decoding phase on heterogeneous NPU platforms (e.g., Ascend 910B) faces severe memory-bound challenges. Th…
cs.LG2026
MR-ImagenTime: Multi-Resolution Time Series Generation through Dual Image Representations
Xianyong Xu, Yuanjun Zuo, Zhihong Huang +4
Time series forecasting is vital across many domains, yet existing models struggle with fixed-length inputs and inadequate multi-scale modeling. We propose MR-CDM, a framework comb…
cs.DC2025
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
Dingyan Zhang, Haotian Wang, Yang Liu +4
Model autoscaling is the key mechanism to achieve serverless model-as-a-service, but it faces a fundamental trade-off between scaling speed and storage/memory usage to cache parame…