3 papers
cs.AI2025
PerPO: Perceptual Preference Optimization via Discriminative Rewarding
Zining Zhu, Liang Zhao, Kangheng Lin +7
This paper presents Perceptual Preference Optimization (PerPO), a perception alignment method aimed at addressing the visual discrimination challenges in generative pre-trained mul…
cs.CV2024
VIP: Versatile Image Outpainting Empowered by Multimodal Large Language Model
Jinze Yang, Haoran Wang, Zining Zhu +3
In this paper, we focus on resolving the problem of image outpainting, which aims to extrapolate the surrounding parts given the center contents of an image. Although recent works…
cs.CV2024
Focus Anywhere for Fine-grained Multi-page Document Understanding
Chenglong Liu, Haoran Wei, Jinyue Chen +7
Modern LVLMs still struggle to achieve fine-grained document understanding, such as OCR/translation/caption for regions of interest to the user, tasks that require the context of t…