6 citations · 6 across the 4 of their papers we have counts for
1 paper · 2 filters
Lingjun Mao, Rodolfo Corona, Xin Liang +2
Most existing vision encoders map images into a fixed-length sequence of tokens, overlooking the fact that different images contain varying amounts of information. For example, a v…