Showing 2024Show all
3 papers · 1 filter
cs.CV2024
Visual Lexicon: Rich Image Features in Language Space
XuDong Wang, Xingyi Zhou, Alireza Fathi +2
We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are of…
cs.CV2024
SegLLM: Multi-round Reasoning Segmentation
XuDong Wang, Shaolun Zhang, Shufan Li +5
We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual…
cs.CV2024
When Does Perceptual Alignment Benefit Vision Representations?
Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler +5
Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range…