5 papers
Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits
Zelong Sun, Jiahui Wu, Ying Ba +2
As social media platforms proliferate, users increasingly demand intuitive ways to create diverse, high-quality portrait collections. In this work, we introduce Portrait Collection…
Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
Weiliang Tang, Dong Jing, Jia-Hui Pan +5
Recent Large Multimodal Models have demonstrated remarkable reasoning capabilities, especially in solving complex mathematical problems and realizing accurate spatial perception. O…
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
Dong Jing, Nanyi Fei, Zhiwu Lu
In the realm of Large Multi-modal Models (LMMs), the instruction quality during the visual instruction tuning stage significantly influences the performance of modality alignment.…
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
Zelong Sun, Dong Jing, Zhiwu Lu
Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images by integrating information from a composed query (reference image and modification text) without training…
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval
Zelong Sun, Dong Jing, Guoxing Yang +2
Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes…