2 papers
cs.RO2026
MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation
Haiyang Sun, Haoxiao Wang, Junming Chen +6
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Action policies are usually evaluated when the current observation largely…
cs.AI2026
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
Haoting Shi, Wenhao Wang, Weicheng Fang +6
Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly through the GUI, often producing inefficient trajectories. Real-world computer w…