Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Fanheng Kong, Jingyuan Zhang, Yahui Liu +8
Multimodal information retrieval (MIR) faces inherent challenges due to the heterogeneity of data sources and the complexity of cross-modal alignment. While previous studies have i…
cs.CV2025
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
Dexian Cai, Xiaocui Yang, Yongkang Liu +4
Existing visual perception systems focus on region-level segmentation in single-turn dialogues, relying on complex and explicit query instructions. Such systems cannot reason at th…
cs.CV2022
A Creative Industry Image Generation Dataset Based on Captions
Xiang Yuejia, Lv Chuanhao, Liu Qingdazhu +3
Most image generation methods are difficult to precisely control the properties of the generated images, such as structure, scale, shape, etc., which limits its large-scale applica…