Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
QK-Normed MLA: QK normalization without full key caching
Yizhou Han, Yao Zhao, Jun Zhou +2
Query-key (QK) normalization stabilizes attention by controlling the scale of queries and keys before the dot product, but is not immediately compatible with Multi-head Latent Atte…
cs.LG2025
Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment
Huanyi Xie, Lijie Hu, Lu Yu +6
In the realm of Text-attributed Graphs (TAGs), traditional graph neural networks (GNNs) often fall short due to the complex textual information associated with each node. Recent me…
cs.LG2025
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
Ling Team, Binwei Zeng, Chao Huang +71
In this technical report, we tackle the challenges of training large-scale Mixture of Experts (MoE) models, focusing on overcoming cost inefficiency and resource limitations preval…