5 papers
Thinking with Anchors: Grounded and Efficient Document Reasoning
Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…
Any2Poster: Any-Source Poster Generation Across Modalities and Domains
Amogh Vinaykumar, Aiden Li, Suozhi Huang +1
Visual posters are a compact medium for communicating dense information, yet progress on automatic poster generation remains difficult to measure because existing evaluations are o…
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
Zihao Lin, Wanrong Zhu, Jiuxiang Gu +8
Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grai…
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
Zihao Lin, Samyadeep Basu, Mohammad Beigi +18
The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for…
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
Binxu Li, Tiankai Yan, Yuanting Pan +8
Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents…