1 paper
Xiaoce Wang, Guibin Zhang, Junzhe Li +3
Existing GUI agent models relying on coordinate-based one-step visual grounding struggle with generalizing to varying input resolutions and aspect ratios. Alternatives introduce co…