most citedExternalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering

6 citations · 6 across the 4 of their papers we have counts for

collaborators

8 papers

cs.SE20266 cited

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering

Chenyu Zhou, Huacan Chai, Wenteng Chen +18

Large language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the…

cs.AI2026

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

Jiachen Zhu, Lingyu Yang, Rong Shan +6

The rise of autonomous GUI agents has triggered adversarial countermeasures from digital platforms, yet existing research prioritizes utility and robustness over the critical dimen…

cs.AI2026

Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation

Weiming Zhang, Jihong Wang, Jiamu Zhou +9

Large Language Models (LLMs) have empowered autonomous agents to handle complex web navigation tasks. While recent studies integrate tree search to enhance long-horizon reasoning,…

cs.LG2026

Adaptive Milestone Reward for GUI Agents

Congmin Zheng, Xiaoyun Mo, Xinbei Ma +10

Reinforcement Learning (RL) has emerged as a mainstream paradigm for training Mobile GUI Agents, yet it struggles with the temporal credit assignment problem inherent in long-horiz…

cs.MA2025

ColorAgent: Building A Robust, Personalized, and Interactive OS Agent

Ning Li, Qiqiang Lin, Zheng Wu +19

With the advancements in hardware, software, and large language model technologies, the interaction between humans and operating systems has evolved from the command-line interface…

cs.CL2025

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Congmin Zheng, Jiachen Zhu, Zhuoying Ou +8

Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final ans…