3 papers
cs.AI2026
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Liya Zhu, Xin Ma, Tao Liu +35
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on re…
cs.CR2026
RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
Di Lu, Bo Zhang, Xiyuan Li +5
Natural-language control offers a promising interface for unmanned aerial vehicles (UAVs), but directly applying self-hosted computer-use agents (SHCUAs) to UAV control introduces…
cs.CR2026
When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents
Di Lu, Yongzhi Liao, Xutong Mu +5
Host-acting agents promise a convenient interaction model in which users specify goals and the system determines how to realize them. We argue that this convenience introduces a di…