works on

From the 2 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CV2026

Thinking with Anchors: Grounded and Efficient Document Reasoning

Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…

cs.CV2026

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Shufan Li, Jiuxiang Gu, Kangning Liu +4

Masked Discrete Diffusion Models (MDMs) have achieved strong performance across a wide range of multimodal tasks, including image understanding, generation, and editing. However, t…

cs.CV2026

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

Shufan Li, Jiuxiang Gu, Kangning Liu +4

Lavida-O is a unified masked diffusion model that combines a lightweight generation branch with a larger understanding branch to perform image understanding, object grounding, imag…

cs.LG2026

FLARE: Diffusion for Hybrid Language Model

Yuchen Zhu, Jing Shi, Chongjian Ge +9

Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck for low-latency deployment. Recent efficien…

cs.GR2026

Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration

Zhuoling Li, Jiarui Zhang, Ping Hu +4

Method illustrations (MIs) play a crucial role in conveying the core ideas of scientific papers, yet their generation remains a labor-intensive process. Here, we take inspiration f…

cs.AI2026

DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation

Zhuoling Li, Hossein Rahmani, Jiarui Zhang +5

The rapid growth of the text-to-image (T2I) community has fostered a thriving online ecosystem of expert models, which are variants of pretrained diffusion models specialized for d…