3 papers
cs.PF2026
Testing the EPYC Conjecture on Real Hardware: MoA-Guided Dense Matrix Multiplication on NCSA Delta (AMD EPYC 7763 Milan)
Lenore M Mullin
A companion empirical study conjectured that MoA-guided dense matrix multiplication would need per-CCD recalibration on AMD EPYC Bergamo; allocation access to that machine was decl…
cs.PF2026
Machine Shape and Hierarchical Blocking: A Mathematics of Arrays Formalization, with an Open Problem in Hierarchical Shape Occupancy
Lenore M Mullin
A companion empirical study found that dense matrix multiplication block sizes calibrated on Apple M1 Pro correspond to two cache-t formulas that mispredict badly on a dierent chip…
cs.LG2026
Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels
Lenore Mullin, Gaetan Hains
The attention mechanism is the dominant computational bottleneck in modern transformer-based AI. Its standard implementation incurs quadratic memory traffic in the sequence length~…