collaborators

9 papers

cs.AI2026

Planning with the Views

Kangrui Wang, Linjie Li, Zhengyuan Yang +7

Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understanding how a single action transf…

cs.CL2026

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

Jiajie Jin, Yuyang Hu, Kai Qiu +15

Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the result…

cs.CV2026

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

Yan Li, Zezi Zeng, Yifan Yang +12

The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for webpage design, offering a flexi…

cs.CV2026

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

Yan Li, Zezi Zeng, Ziwei Zhou +13

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly f…

cs.HC2026

CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning

Minheng Ni, Yutao Fan, Zhengyuan Yang +6

Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However,…

cs.CV2025

Conditional Text-to-Image Generation with Reference Guidance

Taewook Kim, Ze Wang, Zhengyuan Yang +4

Text-to-image diffusion models have demonstrated tremendous success in synthesizing visually stunning images given textual instructions. Despite remarkable progress in creating hig…