1 paper
Zilin Xiao, Ming Gong, Paola Cascante-Bonilla +3
We introduce AutoVER, an Autoregressive model for Visual Entity Recognition. Our model extends an autoregressive Multi-modal Large Language Model by employing retrieval augmented c…