activity
20242026
most citedEnhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing 2025Show all

10 papers · 1 filter

cs.CV2025

MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration

Guangyuan Li, Bo Li, Jinwei Chen +3

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First,…

cs.CV2025

RED: Robust Event-Guided Motion Deblurring with Modality-Specific Disentanglement

Yihong Leng, Siming Zheng, Jinwei Chen +3

Event-guided motion deblurring reconstructs sharp images using the high-temporal-resolution motion cues from event cameras. However, in real capture, thresholding-induced event und…

cs.CV2025

Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior

Zhenning Shi, Zizheng Yan, Yuhang Yu +6

Reference-based Image Super-Resolution (RefSR) aims to restore a low-resolution (LR) image by utilizing the semantic and texture information from an additional reference high-resol…

cs.CV2025

SDMatte: Grafting Diffusion Models for Interactive Matting

Longfei Huang, Yu Liang, Hao Zhang +6

Recent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge r…

cs.CV2025

BokehDiff: Neural Lens Blur with One-Step Diffusion

Chengxuan Zhu, Qingnan Fan, Qi Zhang +4

We introduce BokehDiff, a novel lens blur rendering method that achieves physically accurate and visually appealing outcomes, with the help of generative diffusion prior. Previous…

cs.AI20251 cited

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

Xinbin Yuan, Jian Zhang, Kaixin Li +8

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…