2 papers
cs.CL2026
Block-Sparse Attention with Semantic-Geometric Decoupled Routing
Xinwei Long, Weigao Sun, Weibo Gao +6
Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-…
cs.AR2024
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
Prabhu Vellaisamy, Harideep Nair, Thomas Kang +6
The increasing complexity of deep neural networks (DNNs) poses significant challenges for edge inference deployment due to resource and power constraints of edge devices. Recent wo…