4 papers
AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent
Weidi Luo, Qiming Zhang, Yihao Quan +5
Coding agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn in…
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
Weidi Luo, Tianyu Lu, Qiming Zhang +8
Recent advances in multi-modal large reasoning models (MLRMs) have shown significant ability to interpret complex visual content. While these models enable impressive reasoning cap…
Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
Yanbin Zhang, Hanhui Ye, Yue Bai +4
Workflow automation promises substantial productivity gains in everyday document-related tasks. While prior agentic systems can execute isolated instructions, they struggle with au…
Your Harness is Not Secure: Benchmarking Real-world Threat of Command Line Interface Agent
Weidi Luo, Qiming Zhang, Tianyu Lu +9
Command-line interface (CLI) agents powered by large language models (LLMs) can interpret natural-language requests, plan multi-step tasks, execute shell commands, and modify files…