most citedEvolving Interpretable Constitutions for Multi-Agent Coordination

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

cs.MA20261 cited

Evolving Interpretable Constitutions for Multi-Agent Coordination

Ujwal Kumar, Alice Saito, Hershraj Niranjani +2

Constitutional AI has focused on single-model alignment using fixed principles. However, multi-agent systems create novel alignment challenges through emergent social dynamics. We…

cs.LG2026

Memorization Control in Diffusion Models from Denoising-centric Perspective

Thuy Phuong Vu, Mai Viet Hoang Do, Minhhuy Le +2

Controlling memorization in diffusion models is critical for applications that require generated data to closely match the training distribution. Existing approaches mainly focus o…

cs.AI2025

Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety

Antonio-Gabriel Chacón Menke, Phan Xuan Tan, Eiji Kamioka

Recent work has highlighted the importance of monitoring chain-of-thought reasoning for AI safety; however, current approaches that analyze textual reasoning steps can miss subtle…

cs.CV2025

Region in Context: Text-condition Image editing with Human-like semantic reasoning

Thuy Phuong Vu, Dinh-Cuong Hoang, Minhhuy Le +1

Recent research has made significant progress in localizing and editing image regions based on text. However, most approaches treat these regions in isolation, relying solely on lo…

cs.LG2025

How Effective Is Constitutional AI in Small LLMs? A Study on DeepSeek-R1 and Its Peers

Antonio-Gabriel Chacón Menke, Phan Xuan Tan

Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-cri…