1 paper
Shubham Aggarwal, Lokendra Kumar
The dense output projection in multi head attention scales quadratically with model dimension, contributing significantly to parameter count, memory footprint, and inference cost.…