2 papers
cs.AI2026
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Yikun Fu, Bowen Fu, Zhenyu Wu +10
Computer use agents (CUAs) have advanced rapidly in desktop automation, and a growing number of users deploy CUAs such as OpenClaw on Mac Mini for always-on automation. However, ex…
cs.LG2026
MARTI-MARS: Scaling Multi-Agent Self-Search via Reinforcement Learning for Code Generation
Shijie Wang, Pengfei Li, Yikun Fu +21
While the complex reasoning capability of Large Language Models (LLMs) has attracted significant attention, single-agent systems often encounter inherent performance ceilings in co…