7 papers
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Ping Wu, Haibo Tong, Feifei Zhao +7
Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…
ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
Ximo Zhu, Ruiqi Liu, Rong Wang +8
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local conf…
WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
Jianxun Cui, Ping Wu, Stanisa Peric +2
World models and generative simulators are emerging as interactive testing infrastructure for autonomous driving because they can react to the ego planner and produce counterfactua…
K-Gen: A Multimodal Language-Conditioned Approach for Interpretable Keypoint-Guided Trajectory Generation
Mingxuan Mu, Guo Yang, Lei Chen +2
Generating realistic and diverse trajectories is a critical challenge in autonomous driving simulation. While Large Language Models (LLMs) show promise, existing methods often rely…
ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI
Haibo Tong, Feifei Zhao, Linghao Feng +18
Rapidly evolving AI exhibits increasingly strong autonomy and goal-directed capabilities, accompanied by derivative systemic risks that are more unpredictable, difficult to control…
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
Bing Han, Feifei Zhao, Dongcheng Zhao +4
While fine-tuning services drive the rapid expansion of task capabilities in large language models (LLMs), they are often accompanied by the degradation and reorganization of safet…