1 paper
James Henderson, Alireza Mohammadshahi, Andrei C. Coman +1
We argue that Transformers are essentially graph-to-graph models, with sequences just being a special case. Attention weights are functionally equivalent to graph edges. Our Graph-…