1 paper
Haodong Chang, Hailiang Hu, Zhenrui Wang +5
Attention is a fundamental computational kernel that accounts for the majority of the workload in transformer and LLM computing. Optimizing dataflow is crucial for enhancing both p…