2 papers
cs.LG2025
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
Mark Horton, Tergel Molom-Ochir, Peter Liu +8
Pre-trained transformer models with extended context windows are notoriously expensive to run at scale, often limiting real-world deployment due to their high computational and mem…
cs.AR2021
Neuromorphic Algorithm-hardware Codesign for Temporal Pattern Learning
Haowen Fang, Brady Taylor, Ziru Li +3
Neuromorphic computing and spiking neural networks (SNN) mimic the behavior of biological systems and have drawn interest for their potential to perform cognitive tasks with high e…