1 citations · 1 across the 9 of their papers we have counts for
1 paper · 1 filter
Jingyuan Huang, Zuming Huang, Yucheng Shi +4
Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordina…