1 paper · 1 filter
David M. Chan, Rodolfo Corona, Joonyong Park +3
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representati…