1 paper · 1 filter
Jiayi Zhou, Yang Sheng, Hantao Lou +2
As LLM-based agents increasingly operate in high-stakes domains with real-world consequences, ensuring their behavioral safety becomes paramount. The dominant oversight paradigm, L…