3 papers
cs.LG2026
QK-Normed MLA: QK normalization without full key caching
Yizhou Han, Yao Zhao, Jun Zhou +2
Query-key (QK) normalization stabilizes attention by controlling the scale of queries and keys before the dot product, but is not immediately compatible with Multi-head Latent Atte…
cs.CL2026
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Ang Li, Ben Liu, Bin Han +215
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve,…
cs.CL2023
Self-Evaluation Improves Selective Generation in Large Language Models
Jie Ren, Yao Zhao, Tu Vu +2
Safe deployment of large language models (LLMs) may benefit from a reliable method for assessing their generated content to determine when to abstain or to selectively generate. Wh…