3 citations · 7 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Images are Worth Variable Length of Representations
Lingjun Mao, Rodolfo Corona, Xin Liang +2
Most existing vision encoders map images into a fixed-length sequence of tokens, overlooking the fact that different images contain varying amounts of information. For example, a v…
cs.CV2024
Analyzing The Language of Visual Tokens
David M. Chan, Rodolfo Corona, Joonyong Park +3
With the introduction of transformer-based models for vision and language tasks, such as LLaVA and Chameleon, there has been renewed interest in the discrete tokenized representati…