Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
Seyed Morteza Emadi
Attention scores in transformers are bilinear forms whose maximum magnitude governs overflow risk in low-precision training. We derive a \emp…
cs.LG2026
Exact Attention Sensitivity and the Geometry of Transformer Stability
Seyed Morteza Emadi
We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation. Our main result is the exact identity …