2 papers
cs.LG2026
Efficient Attention via Pre-Scoring: Prioritizing Informative Keys in Transformers
Zhexiang Li, Haoyu Wang, Yutong Bao +1
Efficient attention mechanisms enable long-context transformers but often miss globally important tokens, degrading modeling quality. We introduce a pre-scoring framework that assi…
cs.CL2026
JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
Bingyang Kelvin Liu, Ziyu Patrick Chen, David P. Woodruff
Current autoregressive language models couple high-level reasoning and low-level token generation into a single sequential process, making the reasoning trajectory vulnerable to co…