3 papers
cs.CV2025
On Data Synthesis and Post-training for Visual Abstract Reasoning
Ke Zhu, Yu Wang, Jiangjiang Liu +3
This paper is a pioneering work attempting to address abstract visual reasoning (AVR) problems for large vision-language models (VLMs). We make a common LLaVA-NeXT 7B model capable…
cs.CV2024
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
Yanpeng Sun, Jing Hao, Ke Zhu +6
Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distill…
cs.LG2024
Continual SFT Matches Multimodal RLHF with Negative Supervision
Ke Zhu, Yu Wang, Yanpeng Sun +4
Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiori…