9 papers
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
Zhengyue Zhao, Yingzi Ma, Somesh Jha +3
Large Language Models have shown impressive generative capabilities across diverse tasks, but their safety remains a critical concern. Existing post-training alignment methods, suc…
One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
Yingzi Ma, Zichen Zhu, Ming Jiang +1
On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts. Existing teachers…
MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models
Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu +3
Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autor…
GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving
Yingzi Ma, Chaowei Xiao, Ming Jiang
Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains underexplored, raising unce…
SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
Yingzi Ma, Xiaogeng Liu, Yawen Zheng +1
With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos from a text prompt or an initia…
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
Xiaogeng Liu, Xinyan Wang, Yingzi Ma +2
On-policy self-distillation (OPSD) trains a student on its own rollouts using a privileged teacher, but its standard objective weights all generated tokens equally, implicitly trea…