2 citations · 3 across the 13 of their papers we have counts for
1 paper · 1 filter
Anlin Zheng, Qi Han, Xin Wen +5
In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenize…