Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MOSAIC: Composable Safety Alignment with Modular Control Tokens
Jingyu Peng, Hongyu Chen, Jiancheng Dong +5
Safety alignment in large language models (LLMs) is commonly implemented as a single static policy embedded in model parameters. However, real-world deployments often require conte…
cs.AI2025
Stepwise Reasoning Error Disruption Attack of LLMs
Jingyu Peng, Maolin Wang, Xiangyu Zhao +6
Large language models (LLMs) have made remarkable strides in complex reasoning tasks, but their safety and robustness in reasoning processes remain underexplored. Existing attacks…