collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

PailitaoGR: Latent Think-with-Images for Generative Image Retrieval

Xiaomeng Fan, Yueran Liu, Shengyu Zhou +6

Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradigm to image search, however, is nontrivial…

cs.CV2025

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

Zhanheng Nie, Chenghan Fu, Daoze Zhang +5

Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance…

cs.CV2025

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling

Xianjie Liu, Yiman Hu, Yixiong Zou +3

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding tasks. However, their performance on high-resolution images remains suboptimal. While…

cs.CV2025

Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning

Yukang Lin, Xiang Zhang, Shichang Jia +9

Creative image in advertising is the heart and soul of e-commerce platform. An eye-catching creative image can enhance the shopping experience for users, boosting income for advert…

cs.CV2025

MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding

Daoze Zhang, Chenghan Fu, Zhanheng Nie +7

With the rapid advancement of e-commerce, exploring general representations rather than task-specific ones has attracted increasing research attention. For product understanding, a…