works on

From the 1 of 23 linked papers with an AI index.

activity
20242026
collaborators

23 papers

cs.AI2026

PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks

Hui Xie, Tong Shi, Haotong Qin +3

Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurrent membrane states are commonly retained i…

cs.AI2026

SemPIC: Learning Semantic Position-Independent KV Caches

Hui Xie, Peng Xiao, Yutong Deng +4

The paper introduces SemPIC, a method that learns semantic position‑independent key‑value caches for large language models by training a LoRA‑enabled writer to compile document rep…

cs.CV2026

Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling

Xingyu Zheng, Xianglong Liu, Yifu Ding +4

Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system…

cs.LG2026

An Empirical Study of openPangu Quantization on Ascend NPUs

Tong Shi, Jiacheng Wang, Hui Xie +4

openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive post-training quantization on Ascend NPUs ha…

cs.LG2026

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

Yifu Ding, Jiacheng Wang, Ge Yang +4

Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression method…

cs.CV2026

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models

Jinyang Du, Shenghao Jin, Ziqian Xu +5

Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a large resident parameter footprint…