3 papers
cs.CL2026
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
Xiaoyan Su, Peijie Dong, Zhenheng Tang +8
Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for profess…
cs.CV2026
Enhancing Image Aesthetics with Dual-Conditioned Diffusion Models Guided by Multimodal Perception
Xinyu Nan, Ning Wang, Yuyao Zhai +1
Image aesthetic enhancement aims to perceive aesthetic deficiencies in images and perform corresponding editing operations, which is highly challenging and requires the model to po…
cs.AI2026
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
Yifan Wu, Yiran Peng, Yiyu Chen +12
The performance of autonomous Web GUI agents heavily relies on the quality and quantity of their training data. However, a fundamental bottleneck persists: collecting interaction t…