Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PATH-Bench: Path-Dependent Evaluation of Lifelong Agents
Xidong Yang, Xingyi Zhang, Wenhao Li +7
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely acco…
cs.AI2026
See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch
Xingyi Zhang, Yulei Ye, Kaifeng Huang +2
Block-based programming environments such as Scratch play a central role in low-code education, yet evaluating the capabilities of AI agents to construct programs through Graphical…