2 papers
eess.SP2026
BFLA: Block-Filtered Long-Context Attention Mechanism
Chong Wu, Zhenan Feng, Renjie Xu +5
This paper proposes Block-Filtered Long-Context Attention (BFLA), a training-free sparse prefill attention mechanism for long-context inference. BFLA adopts a two-stage design. In…
cs.DC2025
UniFormer: Unified and Efficient Transformer for Reasoning Across General and Custom Computing
Zhuoheng Ran, Chong Wu, Renjie Xu +2
The success of neural networks such as convolutional neural networks (CNNs) has been largely attributed to their effective and widespread deployment on customised computing platfor…