3 papers
cs.AI2025
Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions
Dong Jing, Nanyi Fei, Zhiwu Lu
In the realm of Large Multi-modal Models (LMMs), the instruction quality during the visual instruction tuning stage significantly influences the performance of modality alignment.…
cs.IR2024
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval
Zelong Sun, Dong Jing, Guoxing Yang +2
Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes…
cs.CV2024
Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts
Jinqiang Long, Yanqi Dai, Guoxing Yang +4
As the research of Multimodal Large Language Models (MLLMs) becomes popular, an advancing MLLM model is typically required to handle various textual and visual tasks (e.g., VQA, De…