6 papers
EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading
Zhihao Wu, Linhai Zhang, Taiyi Wang +4
Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assig…
Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
Zhihao Wu, Gracia Gong, Qinglin Zhu +2
Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access multiple models (today's rea…
AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
Yu Li, Lehui Li, Zhihao Wu +5
Large language model (LLM) agents have demonstrated strong capabilities across diverse domains, yet automated agent design remains a significant challenge. Current automated agent…
SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation
Jingxuan Chen, Derek Yuen, Bin Xie +14
Smartphone agents are increasingly important for helping users control devices efficiently, with (Multimodal) Large Language Model (MLLM)-based approaches emerging as key contender…
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
Taiyi Wang, Zhihao Wu, Jianheng Liu +3
On-device control agents, especially on mobile devices, are responsible for operating mobile devices to fulfill users' requests, enabling seamless and intuitive interactions. Integ…
AppVLM: A Lightweight Vision Language Model for Online App Control
Georgios Papoudakis, Thomas Coste, Zhihao Wu +3
The utilisation of foundation models as smartphone assistants, termed app agents, is a critical research challenge. These agents aim to execute human instructions on smartphones by…