activity
20242026
collaborators

5 papers

cs.CV2026

Thinking with Anchors: Grounded and Efficient Document Reasoning

Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…

cs.CV2026

Any2Poster: Any-Source Poster Generation Across Modalities and Domains

Amogh Vinaykumar, Aiden Li, Suozhi Huang +1

Visual posters are a compact medium for communicating dense information, yet progress on automatic poster generation remains difficult to measure because existing evaluations are o…

cs.CV2026

MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing

Zihao Lin, Wanrong Zhu, Jiuxiang Gu +8

Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grai…

cs.LG2025

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Zihao Lin, Samyadeep Basu, Mohammad Beigi +18

The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for…

cs.CL2024

MMedAgent: Learning to Use Medical Tools with Multi-modal Agent

Binxu Li, Tiankai Yan, Yuanting Pan +8

Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents…