1 paper · 1 filter
Mikhail Galkin, Xinyu Yuan, Hesham Mostafa +2
Foundation models in language and vision have the ability to run inference on any textual and visual inputs thanks to the transferable representations such as a vocabulary of token…