2 papers
cs.AI2026
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
Tong Ye, Kunyang Han, Guozhi Wang +40
Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet industrial deployment faces three persistent gaps. Sandbox training produces a…
cs.CL2026
EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems
Zeyu Liu, Jian Zhong, Rongduo Han +12
Long-horizon interactions with LLM-based assistants require memory systems that preserve and update user states, preferences, and interaction histories. Existing evaluations report…