Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Siting Li, Zhengyang Wang, Simon Shaolei Du +2
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These…
cs.CV2025
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
Siting Li, Xiang Gao, Simon Shaolei Du
While an image is worth more than a thousand words, only a few provide crucial information for a given task and thus should be focused on. In light of this, ideal text-to-image (T2…