2 papers
cs.LG2025
SALS: Sparse Attention in Latent Space for KV cache Compression
Junlin Mu, Hantao Huang, Jihang Zhang +3
Large Language Models capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value cache size and high memory band…
cs.LG2024
DFA-GNN: Forward Learning of Graph Neural Networks by Direct Feedback Alignment
Gongpei Zhao, Tao Wang, Congyan Lang +3
Graph neural networks are recognized for their strong performance across various applications, with the backpropagation algorithm playing a central role in the development of most…