collaborators

6 papers

cs.CV2026

Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions

Bo Zhao, Kairui Guo, Runnan Du +6

Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many seemingly simple cases. We obs…

cs.CV2026

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

Bo Zhao, Maosheng Pang, Chen Zhang +3

Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including vague user prompts, long doc…

cs.CV2026

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

Zeyu Liu, Zanlin Ni, Yang Yue +5

Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely deco…

cs.CV2026

TexEditor: Structure-Preserving Text-Driven Texture Editing

Bo Zhao, Yihang Liu, Chenfeng Zhang +3

Text-guided texture editing aims to modify object appearance while preserving the underlying geometric structure. However, our empirical analysis reveals that even SOTA editing mod…

cs.CV2026

ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation

Xu Zhang, Cheng Da, Huan Yang +3

Existing 1D visual tokenizers for autoregressive (AR) generation largely follow the design principles of language modeling, as they are built directly upon transformers whose prior…

cs.CV2025

SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

Minglei Shi, Haolin Wang, Borui Zhang +11

Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generati…