most citedEnhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

1 citations · 1 across the 4 of their papers we have counts for

collaborators

9 papers

cs.CV2025

SDMatte: Grafting Diffusion Models for Interactive Matting

Longfei Huang, Yu Liang, Hao Zhang +6

Recent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge r…

cs.AI20251 cited

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

Xinbin Yuan, Jian Zhang, Kaixin Li +8

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…

cs.CV2025

Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model

Yang Yang, Siming Zheng, Qirui Yang +6

Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based boke…

cs.CV2025

MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on

Guangyuan Li, Siming Zheng, Hao Zhang +6

Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despit…

cs.CV2025

Photography Perspective Composition: Towards Aesthetic Perspective Recommendation

Lujian Yao, Siming Zheng, Xinbin Yuan +5

Traditional photography composition approaches are dominated by 2D cropping-based methods. However, these methods fall short when scenes contain poorly arranged subjects. Professio…

cs.CV2025

M2N2V2: Multi-Modal Unsupervised and Training-free Interactive Segmentation

Markus Karmann, Peng-Tao Jiang, Bo Li +1

We present Markov Map Nearest Neighbor V2 (M2N2V2), a novel and simple, yet effective approach which leverages depth guidance and attention maps for unsupervised and training-free…