most citedEnhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.AI20251 cited

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

Xinbin Yuan, Jian Zhang, Kaixin Li +8

Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…

cs.CV2024

Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning

Yuti Liu, Shice Liu, Junyuan Gao +4

Image Aesthetic Assessment (IAA) is a vital and intricate task that entails analyzing and assessing an image's aesthetic values, and identifying its highlights and areas for improv…

cs.CV2024

Hero-SR: One-Step Diffusion for Super-Resolution with Human Perception Priors

Jiangang Wang, Qingnan Fan, Qi Zhang +4

Owing to the robust priors of diffusion models, recent approaches have shown promise in addressing real-world super-resolution (Real-SR). However, achieving semantic consistency an…

cs.CV2024

RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution

Jiangang Wang, Qingnan Fan, Jinwei Chen +3

Benefiting from their powerful generative capabilities, pretrained diffusion models have garnered significant attention for real-world image super-resolution (Real-SR). Existing di…

cs.CV2024

Empowering Segmentation Ability to Multi-modal Large Language Models

Yuqi Yang, Peng-Tao Jiang, Jing Wang +4

Multi-modal large language models (MLLMs) can understand image-language prompts and demonstrate impressive reasoning ability. In this paper, we extend MLLMs' output by empowering M…