2 citations · 2 across the 10 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Jingbo Zhou, Yusai Zhao, Qi Bao +12
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents…
cs.AI2026
OmegaUse: Building a General-Purpose GUI Agent for Autonomous Task Execution
Le Zhang, Yixiong Xiao, Xinjiang Lu +12
Graphical User Interface (GUI) agents show great potential for enabling foundation models to complete real-world tasks, revolutionizing human-computer interaction and improving hum…
cs.AI2025
Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
Mengyu Zhang, Siyu Ding, Weichong Yin +2
Reinforcement Learning with Verifiable Rewards(RLVR) has demonstrated great potential in enhancing the reasoning capabilities of large language models (LLMs). However, its success…