2 papers
cs.CV2026
SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
Shanghao Liu, Renze Chen, Size Zheng +4
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptiv…
cs.LG2026
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
Zihao Zheng, Xiuping Cui, Size Zheng +4
As the Mix-of-Experts (MoE) architecture increases the number of parameters in large models, there is an even greater need for model quantization. However, existing quantization me…