7 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.LG2026
When Good Enough Is Optimal: Multiplication-Only Matrix Inversion Approximation for Quantized Gated DeltaNet
Luoming Zhang, Yuwei Ren, Kui Zhang +7
Matrix inversion in chunk-wise parallel linear attention is a major bottleneck for long-context modeling, particularly on NPUs, where forward-substitution-based methods exhibit lim…
eess.AS2023
Speaker Diaphragm Excursion Prediction: deep attention and online adaptation
Yuwei Ren, Matt Zivney, Yin Huang +3
Speaker protection algorithm is to leverage the playback signal properties to prevent over excursion while maintaining maximum loudness, especially for the mobile phone with tiny l…
cs.LG2023★ 7 cited
A Practical Mixed Precision Algorithm for Post-Training Quantization
Nilesh Prasad Pandey, Markus Nagel, Mart van Baalen +3
Neural network quantization is frequently used to optimize model size, latency and power consumption for on-device deployment of neural networks. In many cases, a target bit-width…