18 citations · 23 across the 12 of their papers we have counts for
1 paper · 1 filter
Chi Zhang, Luca Colagrande, Renzo Andri +6
Multi-Head Attention (MHA) is a critical computational kernel in transformer-based AI models. Emerging scalable tile-based accelerator architectures integrate increasing numbers of…