From the 1 of 23 linked papers with an AI index.
23 papers
PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks
Hui Xie, Tong Shi, Haotong Qin +3
Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained i…
SemPIC: Learning Semantic Position-Independent KV Caches
Hui Xie, Peng Xiao, Yutong Deng +4
The paper introduces SemPIC, a method that learns semantic position‑independent key‑value caches for large language models by training a LoRA‑enabled writer to compile document rep…
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
Xingyu Zheng, Xianglong Liu, Yifu Ding +4
Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system…
An Empirical Study of openPangu Quantization on Ascend NPUs
Tong Shi, Jiacheng Wang, Hui Xie +4
openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs ha…
Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression
Yifu Ding, Jiacheng Wang, Ge Yang +4
Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression method…
Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models
Jinyang Du, Shenghao Jin, Ziqian Xu +5
Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a large resident parameter footprint…