3 papers
cs.AI2025
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
Jixiang Hong, Yiran Zhang, Guanzhong Wang +3
Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-model understanding and generation into a single framework. However, LMMs still strugg…
cs.CV2025
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
Feng Ni, Kui Huang, Yao Lu +4
With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for…
cs.CV2025
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
Kui Huang, Xinrong Chen, Wenyu Lv +3
This report introduces PP-DocBee2, an advanced version of the PP-DocBee, designed to enhance multimodal document understanding. Built on a large multimodal model architecture, PP-D…