4 papers · 1 filter
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
Yukang Feng, Jianwen Sun, Chuanhao Li +8
Recent advancements in Large Multimodal Models (LMMs) have significantly improved multimodal understanding and generation. However, these models still struggle to generate tightly…
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
Jianwen Sun, Fanrui Zhang, Yukang Feng +6
Scientific illustrations demand both high information density and post-editability. However, current generative models have two major limitations: Frist, image generation models ou…
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
Jianwen Sun, Yukang Feng, Chuanhao Li +7
Unified multimodal understanding and generation have recently received much attention in the area of vision and language. Existing UniMs are designed to simultaneously learn both m…
VGA: Vision GUI Assistant -- Minimizing Hallucinations through Image-Centric Fine-Tuning
Ziyang Meng, Yu Dai, Zezheng Gong +3
Recent advances in Large Vision-Language Models (LVLMs) have significantly improve performance in image comprehension tasks, such as formatted charts and rich-content images. Yet,…