4 papers · 1 filter
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Mengru Wang, Junfeng Fang, Shuofei Qiao +17
AI models are increasingly used in scientific discovery and human decision-making. Yet how AI models work and what risks they pose remain poorly understood. As AI development becom…
RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
Houcheng Jiang, Boxuan Zhang, Qiyong Zhong +3
Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on promp…
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
Zhepei Hong, Lin Wang, Liting Li +5
Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moderation. Existing turn-level o…
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang +9
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployment…