18 citations · 21 across the 9 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.CY2025
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
Yue Huang, Chujie Gao, Yujun Zhou +5
The Helpful, Honest, and Harmless (HHH) principle is a foundational framework for aligning AI systems with human values. However, existing interpretations of the HHH principle ofte…
cs.MA2025
Multi-Agent Risks from Advanced AI
Lewis Hammond, Alan Chan, Jesse Clifton +41
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These s…