Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
Mark Horton, Tergel Molom-Ochir, Peter Liu +8
Pre-trained transformer models with extended context windows are notoriously expensive to run at scale, often limiting real-world deployment due to their high computational and mem…
cs.LG2024
MonoSparse-CAM: Efficient Tree Model Processing via Monotonicity and Sparsity in CAMs
Tergel Molom-Ochir, Brady Taylor, Hai Li +1
While the tree-based machine learning (TBML) models exhibit superior performance compared to neural networks on tabular data and hold promise for energy-efficient acceleration usin…