1 paper
Kushal Chakrabarti, Nirmal Balachundar
Modern transformer attention is internally multi-agent -- heads compete and coordinate -- yet we train it as if it were a monolithic optimizer. We formalize this gap: cross-entropy…