1 paper · 1 filter
Lin Wang, Junfeng Fang, Dan Zhang +3
The advent of tool-using LLM agents shifts safety monitoring from output moderation to auditing long, noisy interaction trajectories, where risk-critical evidence is sparse-making…