9 papers
When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study
Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or in…
Skill-Guided Continuation Distillation for GUI Agents
Zhimin Fan, Hongwei Yu, Yeqing Shen +9
Improving GUI agents typically relies on behavior cloning on expert trajectories. However, as the current policy deviates from the expert policy, it inevitably encounters policy-in…
WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents
Lu Yan, Xuan Chen, Xiangyu Zhang
LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state. Existing instruction-follow…
AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications
Yifan Sui, Xin Huang, Hongbing Li +14
The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments or open-source applications…
Deep Researcher Agent: An Autonomous Framework for 24/7 Deep Learning Experimentation with Zero-Cost Monitoring
Xiangyue Zhang
We present \textbf{Deep Researcher Agent}, an open-source framework that enables large language model (LLM) agents to autonomously conduct deep learning experiments around the cloc…
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
Xuan Chen, Lu Yan, Ruqi Zhang +1
Large Language Model (LLM) agents increasingly act through external tools, making their safety contingent on tool-call workflows rather than text generation alone. While recent ben…