1 paper
Shawn Young, Xingyu Zeng, Lijian Xu
This paper investigates the fundamental relationship between model capacity and the minimal number of visual tokens required to preserve image semantics. Inspired by the Minimum De…