2 papers
cs.LG2026
When Good Enough Is Optimal: Multiplication-Only Matrix Inversion Approximation for Quantized Gated DeltaNet
Luoming Zhang, Yuwei Ren, Kui Zhang +7
Matrix inversion in chunk-wise parallel linear attention is a major bottleneck for long-context modeling, particularly on NPUs, where forward-substitution-based methods exhibit lim…
cs.DC2025
MoLink: Distributed and Efficient Serving Framework for Large Models
Lewei Jin, Yongqi Chen, Kui Zhang +5
Large language models represent a groundbreaking shift in generative AI. Yet, these advances come with a significant challenge: the high cost of model serving. To mitigate these co…