2 papers
cs.CL2026
LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference
Zhichen Liu, Ruihan Sun, Hengjie Yang +4
Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to sele…
cs.SE2026
GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing
Xiaoyi Chen, Yifei Gao, Yang Xu +3
Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an application and discover defect…