1 paper
Yuqi Zhou, Sunhao Dai, Shuai Wang +3
Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object gr…