2 papers
cs.LG2026
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
Lei Chen, Yuan Meng, Xiaoyu Zhan +2
Large Language Models (LLMs) offer strong capabilities but incur high inference costs due to dense computation and memory access. Training-free activation sparsity is a promising a…
quant-ph2026
Learning to Decode in Parallel: Self-Coordinating Neural Network for Real-Time Quantum Error Correction
Kai Zhang, Zhengzhong Yi, Shaojun Guo +13
Fast, reliable decoders are pivotal components for enabling fault-tolerant quantum computation (FTQC). Neural network decoders like AlphaQubit have demonstrated potential, achievin…