9 papers
ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents
Wenhao Lan, Shan Li, Xinhua Lai +3
Tool-using large language model (LLM) agents read untrusted content, maintain memory, delegate tasks, and invoke tools with external side effects. Terminal attack-success or policy…
HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
Hy Vision Team, Huawen Shen, Zhengyang Tang +20
The paper introduces HyMobileAgent, a vision-native mobile GUI agent that combines large multimodal models with a co-scaling framework for data and environments to enable precise p…
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning
Wenhao Lan, Shan Li, Xinhua Lai +3
Safety alignment requires language models to refuse harmful requests without losing the ability to answer benign ones. Existing robustness evaluations, however, do not reveal wheth…
PhoneBuddy: Training Open Models for Agentic Phone Use
Zhengyang Tang, Xin Lai, Pengyuan Lyu +23
Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matter…
PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions
Chenxin Li, Zhengyao Fang, Zhengyang Tang +18
Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action. However, much of the current mobile-agent literature sti…
PhoneWorld: Scaling Phone-Use Agent Environments
Zhengyang Tang, Yuxuan Liu, Xin Lai +21
A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks…