2 papers
cs.CR2026
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
Yipu Dou, Wang Yang
We study how to allocate a fixed supervised fine-tuning budget when three objectives must be balanced at once: multi-turn safety alignment, low over-refusal on benign boundary quer…
cs.CR2026
AJAR: Adaptive Jailbreak Architecture for Red-teaming
Yipu Dou, Wang Yang
Large language model (LLM) safety evaluation is moving from content moderation to action security as modern systems gain persistent state, tool access, and autonomous control loops…