1 paper
Siddarth Mamidanna, Daking Rai, Ziyu Yao +1
Large language models (LLMs) demonstrate proficiency across numerous computational tasks, yet their inner workings remain unclear. In theory, the combination of causal self-attenti…