4 papers
On Data Synthesis and Post-training for Visual Abstract Reasoning
Ke Zhu, Yu Wang, Jiangjiang Liu +3
This paper is a pioneering work attempting to address abstract visual reasoning (AVR) problems for large vision-language models (VLMs). We make a common LLaVA-NeXT 7B model capable…
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
Huan Liu, Lingyu Xiao, Jiangjiang Liu +4
With the rapid advancement of Multimodal Large Language Models (MLLMs), a variety of benchmarks have been introduced to evaluate their capabilities. While most evaluations have foc…
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
Yanpeng Sun, Jing Hao, Ke Zhu +6
Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distill…
Continual SFT Matches Multimodal RLHF with Negative Supervision
Ke Zhu, Yu Wang, Yanpeng Sun +4
Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiori…